By Tech & Legal Affairs Desk
Updated: September 2026


Main Facts

Newly unsealed court documents from a landmark copyright lawsuit have laid bare internal anxieties, ethical debates, and strategic acknowledgments within two of the most powerful entities in artificial intelligence: Microsoft and OpenAI. The filings, brought to light on Thursday through a high-profile legal battle initiated by The New York Times, reveal that high-ranking employees and researchers at both companies privately questioned the legitimacy of scraping copyrighted journalistic content to train large language models (LLMs).

Among the most explosive revelations is an internal 2023 Microsoft document in which an employee characterized OpenAI’s data-harvesting practices as potentially "the largest theft of labor in human history." The author warned that artificial intelligence models risk setting off a destructive "doom loop" that degrades the very quality of the models themselves by undermining their foundational supply chain.

The lawsuit, originally filed in late 2023 by The New York Times and later joined by eleven other major publishing organizations, contends that OpenAI and its primary financial backer, Microsoft, systematically exploited copyrighted news articles without permission or compensation to build commercially lucrative generative AI tools. As Judge Sidney Stein of the U.S. District Court for the Southern District of New York weighs ongoing summary judgment motions, a steady stream of internal communications, chat logs, and executive depositions are being unsealed, offering the public an unprecedented look behind the curtain of the generative AI boom.


Chronology of the Legal Battle and Internal Warnings

The friction between publishers and AI developers is not a recent development; rather, internal records show that alarms were sounding within Silicon Valley years before the public lawsuits materialized.

  • August 2020: Jack Clark, then-policy director at OpenAI, circulated an internal memo to President Greg Brockman and CEO Sam Altman. Clark warned that the company was developing systems that directly "substitute for the labor of the people that define the ‘culture’ of society." He cautioned that OpenAI risked "becoming the symbol of how Silicon Valley is thoughtlessly stepping into other parts of life and leaving a mess on the carpet." (Clark later left OpenAI to co-found rival AI firm Anthropic).
  • February 2023: An OpenAI engineer noted in internal channels that "no matter how prominently we show the links, users won’t click," directly challenging the narrative pushed by AI companies that chatbots drive meaningful referral traffic back to news publishers.
  • Late 2023: The New York Times filed a formal copyright infringement lawsuit against OpenAI and Microsoft in the Southern District of New York, alleging billions of dollars in statutory and actual damages related to the unauthorized copying and use of millions of its articles.
  • 2023–2024: Throughout this period, internal memos at Microsoft—penned by Brent Hecht, a director of applied science—described large AI models as "a product that destroys its supply chain," predicting that the global public would eventually view the mass harvesting of intellectual property as "an astonishing theft of unprecedented proportions."
  • June 2023 & February 2024: Nick Turley, who led the ChatGPT product team at OpenAI, wrote in internal messages that AI posed an "existential threat" to traditional publishers, explicitly stating that AI products "will get more and more substitutive as they get better" and that they are "largely substitutive, period."
  • Recent Disclosures (September 2026): Judge Sidney Stein ordered the preservation of roughly 20 million ChatGPT conversation logs as part of the evidentiary discovery process. Subsequent batches of unsealed court documents have triggered intense public and legal scrutiny.

Supporting Data and Key Revelations

The unsealed trove of documents offers a stark counter-narrative to the public-facing legal defenses mounted by OpenAI and Microsoft. While corporate legal teams argue that training AI models on public internet data constitutes protected "fair use"—transforming raw news reports into entirely new, derivative works—internal communications reveal a much more cynical or clear-eyed view of the mechanics at play.

Microsoft Staff Asked If AI Scraping Was 'Largest Theft of Labor in Human History'
  1. The Paywall Bypass: In one particularly damning exchange, an OpenAI staff member detailed the creation of a "hack" designed to systematically bypass The New York Times paywall to ingest subscription-locked content. Upon being informed of this workaround, OpenAI President Greg Brockman allegedly replied simply, "ah nice."
  2. The Myth of Referral Traffic: A core argument used by AI firms to appease publishers has been that conversational AI tools act as digital intermediaries, sending curious readers directly to original sources via citations and hyperlinks. However, internal OpenAI engineering logs from early 2023 bluntly concluded that users simply do not click the links, regardless of how prominently they are displayed.
  3. The Scale of Enforcement: The evidentiary burden in this case has proven immense. Judicial orders have forced OpenAI to meticulously preserve and catalog 20 million ChatGPT conversation logs, a testament to the sheer volume of data interrogation required to trace how models generate responses derived from protected journalism.

Official Responses and Corporate Defense

Faced with the public release of these candid internal memos, both Microsoft and OpenAI have scrambled to contextualize or distance themselves from the remarks.

  • Microsoft’s Position: Microsoft formally stated in court filings that the memos written by Brent Hecht do not represent official corporate policy or views. The company argued that Hecht was not a corporate decision-maker and was hired specifically to "present divergent and asymmetric perspectives" as part of his academic and applied science duties.
  • Satya Nadella’s Testimony: Microsoft CEO Satya Nadella testified during depositions that "anything that is paywalled should be licensed by anyone who wants to use it." Nadella further claimed that had he known OpenAI was actively training its models on paywalled New York Times content, he would have exercised Microsoft’s contractual and operational leverage to force OpenAI to retrain its models from scratch. A Microsoft spokesperson later clarified that Nadella was speaking merely to "broad principles" regarding information consumption.
  • Publishers’ Reaction: Attorneys representing the media plaintiffs have seized upon the unsealed records as smoking guns that expose corporate bad faith. Steven Lieberman, an attorney representing the New York Daily News and seven other co-plaintiff publications, remarked: "The world can see what OpenAI and Microsoft thought all along about the fairness of their own behavior."
  • OpenAI and Anthropic: Representatives for OpenAI did not respond to multiple requests for comment from major journalistic outlets regarding the unsealed documents. Anthropic, queried about Jack Clark’s historical 2020 warnings, declined to comment and referred inquiries back to OpenAI.

Broader Implications for the AI Industry and Digital Publishing

The unfolding legal battle between The New York Times and the AI giants represents an existential watershed moment for both the technology sector and the global media landscape. As Judge Stein prepares to rule on crucial summary judgment motions, the fallout from these unsealed documents extends far beyond the courtroom.

1. The Legal Definition of "Fair Use"

For years, the generative AI boom has relied on the legal presumption that scraping the open internet—including news articles, academic papers, and creative writing—falls under the umbrella of copyright law’s "fair use" doctrine. Proponents argue that machine learning models merely "read" text to learn language patterns, much like a human student. However, if the courts determine that these models act as direct commercial substitutes for original journalism—a reality privately acknowledged by OpenAI’s own product leads—the foundational legal architecture supporting modern AI training data could collapse.

2. The Economic Sustainability of Journalism

The internal admission that AI tools siphon audiences away from original publishers without providing compensatory referral traffic validates the core economic grievance of the news industry. If generative AI engines synthesize and deliver breaking news directly to users within a chat interface, readers have little incentive to visit publisher websites, subscribe, or view advertisements. Without subscription revenues or ad impressions, legacy and digital-native newsrooms face a severe funding crisis, threatening the future of original investigative reporting.

3. Corporate Accountability and Internal Dissent

The documents highlight a recurring theme in the tech sector: internal ethical warnings from researchers and policy teams being sidelined in the ferocious, high-stakes race for market dominance. From Jack Clark’s 2020 warnings about leaving a "mess on the carpet" to Brent Hecht’s 2023 predictions of an AI-induced "doom loop," tech companies appear to have recognized the societal and economic externalities of their products long before defending them in court as entirely benign.

As this landmark litigation progresses toward trial or settlement, the transparency forced by judicial discovery is fundamentally altering public and regulatory perception of how generative AI companies acquire their most valuable asset: human creativity and labor.