Skip to main content
ai

Microsoft says virtually nobody was grabbing NYT articles through its chatbot

By the AIdeaFlow Team

Microsoft says virtually nobody was grabbing NYT articles through its chatbot

Microsoft just dropped some numbers in its legal battle with The New York Times and other publishers, and they paint a very different picture than the lawsuit suggests. According to The Verge, the company analyzed 8.2 million Copilot chat logs specifically selected because they contained keywords related to the publishers' websites. These weren't random chats, they were the ones most likely to reproduce copyrighted content.

The results? Only 0.0015% of those conversations reproduced even full sentences from the publishers' work. That's roughly 12,000 instances out of 8.2 million, and Microsoft says substantive reproductions that could actually replace reading the original article were even rarer. The company is using this data to argue that Copilot doesn't function as a substitute for news content, which is central to the copyright infringement claims.

This matters because it reframes the entire debate about AI and copyright. Publishers argue that training on their content and occasionally spitting it back out constitutes theft. Microsoft is countering with hard usage data showing that in practice, their chatbot almost never does what the lawsuits fear most.

The timing is notable too. This comes as multiple publishers have struck licensing deals with AI companies, essentially betting that partnership beats litigation. Microsoft's data suggests that the actual threat to publisher traffic and revenue from chatbots may be far smaller than the legal rhetoric implies.

Of course, the plaintiffs will argue that even rare verbatim reproduction matters when it happens at scale, and that the training itself is the core violation. But if Microsoft can prove its tool genuinely transforms rather than replaces source material in real-world use, it strengthens the fair use defense significantly.

What this means for you: if you're using AI assistants for research, this data confirms what many already suspected. These tools are better at synthesis than substitution. Instead of asking for specific article text, frame requests around analysis and connections. Try this prompt when researching a topic: "I'm researching [topic]. What are the main perspectives and debates? Summarize the key arguments from multiple sources and identify gaps or contradictions I should investigate further." You'll get more original value than trying to extract quotes, and you're using the tool the way it actually works best.

Ready to apply this tech at your business?

Viking Net helps teams in San Antonio and worldwide stay ahead.

Get a Quote