Microsoft exec called AI scraping 'the largest theft of labor in human history'

Newly unredacted filings in The New York Times' copyright lawsuit against OpenAI and Microsoft reveal internal admissions that AI training involved mass scraping, paywall circumvention, and stripping copyright notices. Microsoft data shows its Copilot answer engine cut NYT click-through rates by up to 93%, a 'doom loop' that threatens publishers. OpenAI's own leadership called AI models an 'existential threat' to journalists. The filing details over 91,000 NYT works in OpenAI's mid-training data and 2 million nytimes.com documents from Common Crawl.
It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its 'content supply chain.'
- jacquesm
It's the robbery of all of our culture to sell it back to us at a mark-up. Crimes this large are crimes against humanity. So many people whose life's work got appropriated without consideration, compensation or consent it is baffling.
It is said that at the heart of every great fortune there is a great crime, so it should be no surprise that the most valuable companies on the planet will most likely result from this crime. And given that justice can be bought by those with the most money you can forget about anything coming of this.
- haritha-j
I just don't understand people saying "but a human learning from a book isn't illegal".
How do people not understand that some laws only make sense at a certain scale? One human learning from resources and being added to the labour pool is not the same as an infinitely copyable entity doing the same thing. One has negligible impact on the demand for the original, and the other replaces 99% of the demand."
And creating a rule that says you cannot train on any material unless the rights holder authorises it via license is not complicated. That will creat a amrketplace where creators can decide the price for their content. It's just inconvenient.
- 47282847
“Information wants to be free“.
It’s not “theft of labor”; the work was already done. If anything it is theft of “intellectual property” (aka “copyright infringement”), if you believe that is a thing, but not of the “labor” that went into it.
My personal take: anyone producing content, everyone’s creativity, is fed by something that others did before. We’re all standing on the shoulders of giants composed of previous generations and their “content’s” distribution and dissemination. I have an immense gratitude for all the labor before me that I was and am allowed to partake; without that, I would be nothing. Sharing information is an act of love; gatekeeping it is short-sighted greed. New technologies have always “killed” previous “labor”, out of which new opportunity grows. I just wished the collected data was public. I hope we all get a mega-leak at some point.
- heaney-555
LLMs are not compression algorithms. From an information theory perspective, that's impossible given their size.
Thus, a distinction needs to be made between viewing material to _learn_ and viewing material to _verbatim repeat_.
It's not illegal to read the New York Times and then start giving paid advice based on what you learned, as long as you don't repeat the text verbatim.
- sajithdilshan
If someone asked what is 'the largest theft of labor in human history' I would have thought slavery.
- Kuyawa
Google has been scraping everything from us since day one. Meta, Microsoft, Github, Slack, Reddit, StackOverflow, big and small, every single app that interacts with people uses our own data to make money and create walled gardens. I haven't seen a single one opening their silos to the world. That's our data, we produced it, you captured it and now you think it's yours
So no, your cries for regulating others because you are losing the race won't work this time.
- cmiles8
The evidence here is quite damning for OpenAI and Microsoft is clearly trying to distance themselves from OpenAI’s behavior here.
- TutleCpt
The most shocking point is that they have a Microsoft exec who knows what he's talking about.
- juvvel
I wouldn't have a problem with working off the fruits of other people's labor because most of us are essentially doing that everyday anyway, the issue is that big tech companies (want to) reap all the benefit and create profit from something that should be accessible to everyone. Everything is getting privatized -- housing, water, electricity, and now, thinking and knowledge. We are heading towards a world where you have to pay even more excessive fees just for existing and for completing any basic task.
- tom2026hn
The problem isn’t just “stealing the fruits of human labor”, it’s also driving down the value of human skills and even taking away human jobs.