NYT v. OpenAI: what creative teams should actually do
The training-data lawsuit that gets the most attention is also the most misunderstood. Here is what creative teams using OpenAI models should and should not change about their workflow.
The New York Times lawsuit against OpenAI is reshaping how lawyers and technologists think about AI training data, but it has been widely misread by downstream users. Most creative teams using OpenAI models do not need to stop what they're doing; they need clearer boundaries, better contracts, and better provenance records.
What the case is actually about
In December 2023, The New York Times Company sued OpenAI and Microsoft in the Southern District of New York, alleging that millions of Times articles were copied without permission to train OpenAI models, and that those models can sometimes emit outputs that closely track Times content. The complaint asserts copyright infringement, DMCA violations, and related claims tied to both training and specific output snippets that appear substantially similar to NYT articles.
In April 2025, Judge Sidney Stein largely refused to dismiss the core copyright claims, allowing the Times to proceed on direct and contributory infringement theories while trimming certain unfair-competition and DMCA counts. The case is active and fact-intensive: the court has also ordered OpenAI to preserve large volumes of chat logs and output data for discovery, which is driving a parallel fight over data retention and privacy.
What it does not change for downstream users today
For most creative teams using ChatGPT or the OpenAI API, nothing in NYT v. OpenAI has suddenly made your day-to-day workflows unlawful or categorically unsafe, as long as you follow a few existing ground rules.
- Training-data liability is upstream. The core dispute is about OpenAI's use of NYT content in training and about model behavior, not about ordinary end-users relying on non-verbatim, transformed outputs in their work.
- You contractually own API outputs, but not risk-free. OpenAI's terms assign you ownership of API outputs and allow commercial use, while capping OpenAI's liability and disclaiming any warranty that outputs are non-infringing. That means you get contractual rights but still need your own clearance practices for high-risk use cases. (For a comparison of how other vendors handle this, see WCR Legal's breakdown.)
- Enterprise tiers include some copyright shield; free tiers don't. OpenAI's "Copyright Shield" program covers certain IP claims for ChatGPT Enterprise and API customers, but not for free ChatGPT tiers, and it comes with carve-outs and a liability cap. If you're on an enterprise plan, this is a meaningful buffer; if you're not, assume you're exposed like with any other vendor tool.
- Output-use liability is still context-specific. Using AI to draft internal copy or first drafts is not the same as bulk reproducing paywalled articles or logos; most creative work today is low-risk when human teams edit, localize, and blend AI material with their own assets. Existing copyright law already treats those differently.
- Your own contracts matter more than the lawsuit. Agency MSAs, client production agreements, and your terms of use can allocate who bears IP risk for AI-assisted work, regardless of what happens between NYT and OpenAI. Many teams already push residual risk to clients or carry E&O insurance for that tail.
In other words: for "normal" use (drafting copy, brainstorming, summarizing internal material, generating bespoke images that you then edit), NYT v. OpenAI has not created a new, special category of danger overnight. It has just raised the stakes for OpenAI and sharpened the questions sophisticated buyers already ask.
What it might change later
Where this case could touch creative teams is indirect and medium-term. The main scenarios to watch are:
- Model behavior and training sources might shift. A ruling that is hostile to certain training practices could push providers toward more licensed, curated, or first-party corpora, which would change how models sound and what they're allowed to emit. You'd see this as more conservative refusals, different style, or higher pricing for "fully cleared" tiers.
- Enterprise contracts may harden. Large customers are already negotiating stronger warranties, indemnities, and data-retention rules; a precedent here could make that standard, and small teams may see stiffer terms or clearer exclusions around "news-like" outputs and scraping.
- Discovery expectations will rise. The preservation orders in this and related litigation show courts are willing to treat AI logs and outputs as discoverable evidence and to demand retention at scale. That flows downstream: if you ship AI-assisted work into regulated or high-value environments, expect future disputes to ask "what prompts and outputs did you rely on, and when?"
- Certain usage patterns may become red flags. If a court eventually articulates examples of outputs that are too close to copyrighted sources, those patterns (e.g., verbatim or near-verbatim reproduction of paywalled reporting) may become clearly off-limits, even for end-users.
None of this is settled law, and none of it is legal advice. But all of it argues for better provenance and governance on the creative team side: knowing which model you used, on what settings, for which project, and what exactly you shipped.
What it does not mean you should do
Equally important is what you shouldn't overreact to:
- You do not need to ban OpenAI outputs from early-stage ideation, internal drafts, or pitch work.
- You do not need to move everything off OpenAI just because NYT sued; the key issues will apply, in different flavors, to any general-purpose model trained on broad internet data.
- You do not need to pretend AI wasn't used. Some provider terms now explicitly bar misrepresenting AI outputs as human-generated, and regulators are moving in the same direction.
The smarter move is to contain and document AI use, not to pretend it never happened.
What we recommend in Archibal
For Archibal users, the right response to NYT v. OpenAI is not to change everything - it's to tighten the paper trail around what you already do:
- Track the OpenAI tier per project. Record, in the project metadata, whether the work used OpenAI Free/Plus, API, or an enterprise plan, and which model class. That matters for indemnity and for how courts might view your diligence.
- Capture prompts and outputs with C2PA-style hashes where possible. Use Archibal's integration to hash prompts and critical outputs and embed or associate them in a C2PA manifest, so you can later prove "this is what we actually saw and used on this date." You don't need to log every keystroke - focus on material that shipped or heavily shaped the final work.
- Run sign-off through Archibal, not email. Treat AI-assisted work like any other risky asset: route it through a formal review and sign-off workflow so there is a clear lawyer-of-record or approver tied to a timestamped certificate. That gives you a defensible story later: which version was approved, under which assumptions, by whom.
- Separate low-risk and high-risk uses. Use OpenAI freely for internal ideation and structure, but require an Archibal project and sign-off for anything that will be published, heavily monetized, or used in regulated contexts. The more exposed the asset, the more you want provenance, auditability, and counsel sign-off.
- Keep your contracts up to date. Make sure your client MSAs and SOWs say that AI-assisted methods may be used, clarify who owns outputs, and allocate residual IP risk realistically (for example, "standard industry clearance, no strict guarantee of non-infringement"). Archibal's certificates and logs back up those assurances in practice.
Put bluntly: NYT v. OpenAI is a wake-up call for vendors and governance, not a reason for creative teams to abandon AI. Keep using the tools - just leave a better trail of what you did, and make sure someone senior owns the decision to ship.
If you produce AI work for clients, see how Archibal works for creators and agencies.