The lawsuit filed by the Seattle Times and Newsday against OpenAI and Microsoft raises an increasingly important question for the digital economy: how far can artificial intelligence companies go in using news content to develop their products without authorization or compensation for the organizations that produce it?

Two American newspapers take the AI dispute to court

The Seattle Times and Newsday, two newspapers in the United States, have filed a lawsuit against OpenAI and Microsoft, accusing the companies of copyright infringement over their use of news content in the development of artificial intelligence systems.

The lawsuit was filed in federal court in Manhattan and claims that the companies collected content from the newspapers’ websites, including paywalled material, and used the articles in datasets connected to the training and operation of AI products.

The allegations put a question at the center of the dispute that goes beyond the two publications: can content produced by news organizations be used to build AI products that later compete for the same audience attention?

What the newspapers allege

According to the lawsuit, the systems have been capable of reproducing or paraphrasing content from the newspapers, in some cases with a high degree of similarity to the original material.

The publishers argue that this type of behavior could reduce the economic incentive for readers to visit their websites directly, subscribe to their services or consume their reporting at the original source.

The lawsuit seeks financial compensation as well as measures related to copies of the content and models or datasets that, according to the plaintiffs, incorporated copyrighted material.

Paywalled content is one of the central issues

Editorial representation of news articles behind a paywall being incorporated into artificial intelligence systems

The lawsuit challenges the use of news content, including paywalled material, in the development of AI systems.

One of the most sensitive elements of the case is the allegation that the content used was not limited to information freely available on the internet. The newspapers claim that paywalled material was also collected.

That makes the economic debate more complicated. A paywall is not simply designed to restrict access. It is part of a business model that turns journalism into revenue through subscriptions.

When AI answers instead of the website

The problem raised by the newspapers is not limited to copying during training. There is also a dispute over what happens after the model is deployed.

If a user can obtain a chatbot response that sufficiently reproduces or summarizes a news report, part of the value that once depended on a visit to the website can instead be captured by the AI interface.

That shift is precisely what makes the case relevant to publishers. The content is still produced by a news organization, but the relationship with the reader can increasingly take place inside another platform.

The dispute involves training and operation

According to the allegations presented in the lawsuit, the articles were incorporated into datasets used to train and operate products such as ChatGPT, Microsoft Copilot and AI features in Bing.

The legal question, however, is not settled simply because the content was found on the internet. The central issue will be determining which uses are permitted under U.S. copyright law and which cross the boundaries of that protection.

OpenAI and Microsoft face a question bigger than this lawsuit

The case comes at a particularly sensitive time for OpenAI and Microsoft. The companies are already facing other disputes over the use of copyrighted works in AI training.

The New York Times lawsuit, filed in 2023, has become one of the most prominent legal battles over the issue. The new lawsuit now adds two more news organizations to the dispute and increases pressure on the way generative AI systems are developed.

The companies’ position

Microsoft said it was surprised by the lawsuit and indicated that it was willing to engage in discussions over the dispute.

The position advanced by technology companies in cases of this kind is that training AI models may constitute a transformative use of publicly available information and therefore fall under the concept of fair use, which is recognized under U.S. copyright law.

That does not mean the issue has been settled. Fair use is a legal analysis that depends on the circumstances of each case, and courts still need to establish clearer boundaries for training generative AI models.

A dispute that has already reached the political debate

The issue gained even more weight this week after the U.S. government presented arguments supporting OpenAI’s position in its legal dispute with the New York Times.

The government argued that restricting certain forms of AI training could harm innovation and U.S. competitiveness in artificial intelligence.

At the same time, publishers argue that allowing unrestricted use of copyrighted content could transfer economic value from the creative industries to technology companies without providing proportional compensation to the organizations that produce the content.

What the lawsuit could change for AI companies and publishers

Representation of a balance between copyright protection and the development of artificial intelligence models

The dispute puts two economic interests in conflict: the need for data to develop AI and the protection of content that generates revenue for its producers.

The outcome of the lawsuit is still far from certain, but the dispute already highlights a strategic problem for the industry.

AI companies need vast amounts of information to develop increasingly capable models. Publishers, on the other hand, depend on exclusivity, audience engagement and their ability to turn original content into revenue.

Content licensing is gaining ground as an alternative

One possible response to the conflict is content licensing.

The industry already has agreements between AI companies and media organizations. The question is whether this model will be enough to create a sustainable economic relationship between the two sides or whether courts will establish different rules for content used in training.

For companies that rely on proprietary information, the definition of these boundaries could affect costs, data availability and product development strategies.

The precedent matters to far more than major newspapers

The potential impact is not limited to large American newspapers.

Blogs, specialized websites, independent publishers and other digital content producers also need to consider how their content will be used by systems capable of collecting, synthesizing and answering questions based on information available online.

The international debate already shows how the issue of training AI with copyrighted content has moved beyond a dispute between technology companies and major news organizations. Governments, regulators, creators and companies that depend on content to generate revenue are now part of the discussion.

The battle over content could reshape the economics of the internet

The dispute involving the Seattle Times, Newsday, OpenAI and Microsoft comes at a time when AI systems are taking on an increasingly important role in how people discover information.

A user who once searched for a question, visited several websites and compared answers can increasingly receive a synthesized response directly through an AI interface.

That creates an important economic shift: content can continue to be produced outside an AI platform, while the interaction with that content and part of the value it generates take place inside the platform.

The risk for those producing information

For a publisher, fewer visits can mean less advertising revenue, fewer subscriptions and a reduced ability to finance new reporting.

That is one reason the copyright debate cannot be viewed solely as a technical dispute over datasets. It also involves who ultimately finances the production of the information that feeds artificial intelligence systems.

The debate connects with other recent discussions over the use of copyrighted content in model training. Notícia Tech has already examined how governments and companies are debating the limits of this practice internationally in an analysis of AI training with copyrighted content.

The next chapter will be decided in court and in the market

Representation of a newsroom facing an artificial intelligence interface, symbolizing the future of news production and distribution

The conflict between publishers and AI companies could help determine how content and artificial intelligence will be monetized in the next phase of the internet.

The lawsuit filed by the Seattle Times and Newsday alone cannot determine the future of copyright in artificial intelligence. But it shows that the relationship between content producers and model developers is entering a more contentious phase.

For AI companies, the question is how to continue developing competitive models in the face of potential restrictions on access to certain types of content. For publishers, the challenge is preventing their investments in information from becoming raw material for products that could reduce the economic value of the original source.

What to watch next

The next legal developments will be important in determining how U.S. courts interpret the use of copyrighted content in the training and operation of generative AI systems.

It will also be important to watch whether more news organizations choose litigation or whether the industry moves toward licensing agreements, with clearer rules governing access, training, attribution and compensation.

The central issue is bigger than OpenAI or Microsoft: if artificial intelligence is becoming a new layer through which people access information, the economics of that information will also need a new way to distribute value between those who produce the content and those who build the technology that uses it.