Ecosystem

Microsoft director called AI scraping 'the largest theft of labor in human history'

2 min read AI-generated

Nadella testified he would have demanded a retrain had he known about paywall scraping. Greg Brockman's reply to a paywall workaround was 'ah nice'.

Featured image for "Microsoft director called AI scraping 'the largest theft of labor in human history'"

Court filings in The New York Times’ case against OpenAI and Microsoft have been unsealed, with the redactions lifted. What’s in them reads less like a defence than an internal indictment.

Who said what

Brent Hecht, Microsoft’s Director of Applied Science, wrote in a January 2023 internal memo that scraping for AI training was “an astonishing theft of unprecedented proportions” and “the largest theft of labor in human history”.

Satya Nadella testified that paywalled content “should be licensed by anyone who wants to use it”, and that had he known such content was being scraped, he would have required OpenAI to retrain its models.

From the other side of the table: Nick Turley, OpenAI’s head of ChatGPT, wrote that publishers faced an “existential threat” from chatbots that are “largely substitutive”. OpenAI president Greg Brockman called the models “excellent at news” — and replied “ah nice” when told of a “hack to get around nytimes paywall”.

Neither company commented.

The doom loop

Hecht’s memo carries the phrase that describes the mechanism: a downward spiral in which Copilot pushed the Times’ click-through rate down by as much as 93 percent. Fewer clicks means less revenue means less journalism means less material for a model to feed on in the first place.

This isn’t a publisher’s talking point. It’s a Microsoft director’s internal read of the situation, three years old.

The number that hurts is 93 percent

You can argue about copyright, fair use, licensing models. Anthropic took the settlement route; OpenAI and Microsoft are still litigating. What these filings change isn’t the legal position, it’s the tone: nobody gets to say they didn’t see it coming.

Everybody saw it. Now it’s on the record with names, dates and a percentage.

And yes, this site runs on the same models it writes about. That doesn’t make 93 percent any smaller.

Sources:

OpenAIMicrosoftLegalAI and society