Microsoft director calls AI scraping 'the largest theft of labor in human history'
Microsoft director: AI scraping 'the largest theft of labor in human history'

Leaked legal briefs from the New York Times' copyright lawsuit against OpenAI and Microsoft reveal damning internal statements. A Microsoft director labeled AI scraping 'the largest theft of labor in human history,' while an OpenAI head called ChatGPT an 'existential threat' to publishers. The documents, sealed at the companies' request, were filed in support of the NYT's motion for summary judgment.
AI scraping is the largest theft of labor in human history.
- adamddev1
I wrote something original in a language learning grammar online. I coined a term to describe something about how a certain language with negative language functions.
I tried talking to an LLM about it and sure and enough, it had scraped it and learned to use that terminology and explanation that I made up. I asked the LLM, "Where did that concept/term come from, who made it up?" It did a bunch of searching and tried to site a whole bunch of other sources which did NOT contain the terms of the explanation I had written. It simply would not cite or mention my source. It kept parroting my material semi-correctly and hiding the source. By any other actor that would be an egregious act of sloppiness, dishonesty, and plagiarism.
Generative AI is speed-running a widespread corruption of truth. I do not believe any of the gains are worth this.
- buzer
Discussion from yesterday (811 comments): https://news.ycombinator.com/item?id=49752056
- eschaton
Pretty rich statement in a world where slavery exists.
- muchdoubt
Only to be topped by the massive amount of IP that Copilot is possibly secretly harvesting at present. We already believe OpenAI/Anthropic are doing it, what’s to stop Microsoft from attempting to improve their competitive advantage by surreptitiously using their role as a MiTM between end user Corporations and Model servers.
- fxtentacle
if he “had been made aware that OpenAI has scraped and trained on information that was behind a paywall,” the company would have required OpenAI “to retrain its models.”
... yeah, sure.