Policy

Microsoft Pushes for Summary Judgment in Publisher Lawsuit, Citing Rare Copilot Text Matches

Filings in the consolidated copyright case show less than 1 percent of surveyed chat logs reproduced short excerpts of news text.

  • Microsoft is urging a federal judge to dismiss copyright infringement lawsuits brought by major news publishers and authors, arguing that its Copilot assistant virtually never outputs verbatim exce…
  • According to new court filings detailed by The Verge, Microsoft submitted 8.2 million Copilot chat logs during discovery to an expert retained by the news plaintiffs.
  • Results were similarly sparse across other plaintiff groups in the consolidated proceedings.
Microsoft Pushes for Summary Judgment in Publisher Lawsuit, Citing Rare Copilot Text MatchesThe Scale Report

Microsoft is urging a federal judge to dismiss copyright infringement lawsuits brought by major news publishers and authors, arguing that its Copilot assistant virtually never outputs verbatim excerpts from copyrighted materials.

According to new court filings detailed by The Verge, Microsoft submitted 8.2 million Copilot chat logs during discovery to an expert retained by the news plaintiffs. The logs were filtered specifically for keywords associated with the publishers' domains. Across that dataset, only 59,545 conversations, or less than 1 percent, included at least 16 words matching the news content used to ground the model.

Results were similarly sparse across other plaintiff groups in the consolidated proceedings. An expert representing the Center for Investigative Reporting identified 51 instances of substantial overlap with the outlet's investigative reporting. Meanwhile, an expert for the suing authors discovered only 24 responses containing at least 30 identical words across the 8.2 million logs, with matching text detected in only 10 of the 212 evaluated books.

Microsoft contends these metrics disprove the publishers' claims that generative chatbots function as market replacements for original journalism and books. In its motion for summary judgment, the company stated that the isolated reproduction of short phrases does not diminish the transformative nature of training large language models under fair use doctrine.

The plaintiffs rejected Microsoft's interpretation of the discovery evidence. Ian Crosby, lead counsel for The New York Times, maintained that the records show intentional infringement. Crosby stated that the documents and testimony uncovered during discovery lead to only one conclusion: Microsoft and OpenAI stole from The New York Times to make commercial products that substitute for its journalism, threaten its business, and undermine its industry, adding that the publication looks forward to holding both companies accountable.

Why It Matters

The dispute serves as a crucial legal benchmark for the generative artificial intelligence industry. If the court grants summary judgment on fair use grounds, it will provide substantial legal cover for technology companies scraping internet content for training datasets without licensing deals. Conversely, a ruling that sends the dispute to trial would elevate financial and operational risks for foundational model developers.

The consolidated litigation includes claims against both Microsoft and partner OpenAI. The legal fight recently drew involvement from the Trump administration, which submitted a statement of interest backing OpenAI in the proceedings. If the court denies Microsoft's motion, the case will advance toward a full trial.

Reporting based on coverage from AI | The Verge.

The daily brief

The biggest stories in AI, venture, sports business and culture - once a day.

One short email from The Scale Report. No spam, unsubscribe any time.

Read next