Ads_970x250

Why India's AI Training Ruling Leaves Sourcing Unsettled

The Delhi High Court's preliminary fair dealing ruling and Anthropic's $1.5 billion US settlement show why model training and data acquisition need to be assessed separately.

Topics

  • Key Takeaways

    01

    India’s first substantive ruling on AI training and copyright provides developers with a preliminary fair-dealing defense, not a general right to use copyrighted material.

    02

    Anthropic’s settlement shows how a favorable finding on model training may coexist with substantial exposure over how source material was obtained and retained.

    03

    Companies should preserve the source, rights and version history of important datasets before a dispute, transaction or regulatory inquiry makes those records necessary.

    On July 24, the Delhi High Court refused ANI Media’s request for an interim injunction against OpenAI. Justice Amit Bansal found, on a preliminary basis, that storing and using ANI’s articles to train the models behind ChatGPT fell within the fair dealing protection for research under Section 52(1)(a) of the Copyright Act. It was India’s first substantive judicial ruling on AI training and copyright.

    Four days earlier, a US court had approved Anthropic’s $1.5 billion settlement with authors who accused the company of building part of its book library from pirated copies. The settlement followed a ruling that deemed the use of books to train Claude fair use, but left Anthropic exposed for the acquisition and permanent retention of millions of pirated books.

    Against that backdrop, the Indian order gives developers a serious legal argument. Research, the court held, is not limited to work performed directly by humans, and a commercial purpose does not by itself defeat the defense. ANI had also failed at this stage to show that ChatGPT memorized or substantially reproduced its articles.

    On sourcing, the court rejected ANI’s argument that the “non-infringing copy” language in the explanation to Section 52(1)(a) applied to all works stored electronically. Yet it also recorded that ANI had not accused OpenAI of using an unauthorized source or breaking through its paywall. The articles were freely available on ANI’s website. The record therefore did not require the court to decide the consequences of piracy, of defeated access controls, or of a supplier licensing material it did not own. Companies still need to know where their training data came from. 

    The Court Limited the Lawful Copy Requirement 

    India’s Copyright Act contains no express exception for AI training or text and data mining. Developers must bring their conduct within provisions written for other uses.

    “There is presently no specific statutory exception under Indian copyright law for the use of copyrighted works to train AI models,” said Tanisha Khanna, a partner at Trace Law Partners.

    ANI argued that the explanation to Section 52(1)(a) protected electronic storage only when the material was not itself an infringing copy. The judge disagreed. From the provision’s wording and punctuation, and its references to “lawful” copies elsewhere, he concluded that the limitation concerned computer programs rather than literary works stored electronically.

    That meant ANI could not defeat OpenAI’s fair dealing defense simply by calling the temporarily stored articles infringing copies. The judgment then gave a factual reason: ANI had not alleged an unauthorized source or paywall circumvention. Because the articles were freely available on ANI’s website, the court said OpenAI had not obtained an infringing copy “in any event.”

    The finding does not establish that acquisition never matters. The court interpreted the explanation to one fair dealing provision on facts that did not involve pirated or access-controlled material. Other copyright provisions, contracts, technological protection measures or the removal of rights-management information may raise separate issues.

    ANI’s sample articles also postdated the April 2022 training cutoff for GPT-4 and the April 2024 cutoff for GPT-4o. ANI supplied no other examples showing use in training. The court still considered fair dealing because OpenAI admitted temporarily storing ANI works during training.

    The order is expressly provisional and says its observations will not affect the outcome. Dinesh Jotwani, co-managing partner at Jotwani Associates, cautioned against treating it as a universal exemption. Courts may still consider whether use is confined to training, whether outputs reproduce protected expression, and whether the system harms the owner’s market, he said.

    “There is presently no specific statutory exception under Indian copyright law for the use of copyrighted works to train AI models.”

    — Tanisha Khanna, Partner, Trace Law Partners

    The Anthropic Case Separates Training From Acquisition

    In June 2025, US District Judge William Alsup held that Anthropic’s use of books to train Claude was fair use. He separately found that scanning purchased print books to create digital library copies was fair use. But downloading millions of books from pirate sites to build a permanent general-purpose library was not protected, even though Anthropic later used some of those books for training.

    Anthropic settled the remaining claims before trial. On July 20, 2026, US District Judge Araceli Martínez-Olguín granted final approval to the $1.5 billion settlement. It covers 482,460 works, with roughly $3,000 allocated to each eligible work. Reuters described it as the largest known settlement in a US copyright case.

    The settlement resolved rather than decided the remaining claims, and the US ruling is not binding in India. It remains relevant because the Delhi court cited Bartz and the US court assessed training, purchased books and pirated library copies separately.

    Recent complaints are testing similar distinctions. Sony Music Publishing and Warner Chappell sued Anthropic in August over copyrighted lyrics and sheet music allegedly obtained through torrents and other sources. WikiHow sued OpenAI over 11,211 articles. These are allegations, not judicial findings.

    India’s policy debate could produce an explicit source requirement. In December 2025, the Department for Promotion of Industry and Internal Trade published Part I of its Working Paper on Generative AI and Copyright. It proposes a mandatory blanket license for lawfully accessed protected works, with royalties payable after commercialization.

    Developers could not use the proposed license to bypass paywalls or technological protection measures. Once lawful access and any required payment were secured, they could train without further permission. “The rightsholders will not have the option to withhold their works for use in the training of AI Systems,” the paper says.

    The consultation closed in February 2026, and the proposal has not been enacted. MeitY’s India AI Governance Guidelines also identified copyrighted training material as a legal gap requiring legislative attention.

    Provenance Is Becoming a Management Control

    The sourcing question extends beyond businesses training foundation models. Indian companies also fine-tune open models, adapt third-party systems, and combine customer records with outside material.

    Nasscom’s India Generative AI Startup Landscape 2025 counted more than 890 active generative AI startups in the first half of 2025, up from more than 240 a year earlier. Among the startups surveyed, 78.8% used proprietary customer data for training or fine-tuning, 66.7% used public-domain datasets, 45.5% synthetic data, 39.4% open government datasets and 33.3% purchased third-party datasets. 

    CoRover, the developer of BharatGPT, says its models use licensed or permissioned material, proprietary and domain-specific information, publicly available content, and synthetic data in proportions that vary by model and training stage.

    “Public availability does not, by itself, establish that content is free of copyright or unrestricted for use,” said founder and CEO Ankush Sabharwal.

    CoRover says it reviews sources, licenses, permissions, contracts, and intended uses before incorporating material. The company says it can investigate whether content appeared in a training corpus, model version or retrieval knowledge base when its records permit tracing. These are company representations rather than an independent audit.

    A useful record should show where material came from, when and how it was collected, and what terms applied. It should identify the rights or statutory exceptions relied on, preserve relevant supplier assurances, and link dataset versions to the models and fine-tuning runs that used them.

    Khanna and Jotwani also recommended testing for memorization and substantial reproduction. The Delhi court considered both training use and whether ChatGPT returned substantially similar material or competed with ANI’s content.

    The European Union already requires providers to disclose some of this information. Providers of general-purpose AI models that were placed on the market since August 2, 2025, must publish a training-content summary. Enforcement began on August 2, 2026, while providers of earlier models have until August 2, 2027. The template covers major datasets, licensed sources, scraped data and prominent web domains rather than every work. 

    An undocumented dataset can complicate legal reviews, customer inquiries and due diligence for investments or acquisitions. Management can respond faster if it knows what entered a system and on what basis.

    Implications for Leaders

    Senior Executives

    Identify products that depend on externally sourced data and decide which sources require closer review. Do not treat public availability, payment to a vendor, and permission from a rights holder as interchangeable. Record the basis on which each material source is being used.

    Product, Data and Legal Teams

    Maintain a dataset register covering source, collection method, date, applicable terms, supplier assurances, and the model versions that used the material. Add proportionate tests for memorization and substantial similarity to release reviews, and retain the results.

    Boards and Risk Committees

    Ask how quickly the company could answer a claim involving one named work. The response should identify the source, contractual position, affected models, testing history, and available mitigation. Continue to track the ANI suit and DPIIT proposal rather than treating the interim order or working paper as settled law.

    The ANI order gives Indian developers room to argue that model training is research under Section 52. It also rejects one attempt to impose a non-infringing-copy condition on electronically stored literary works. But the court’s separate finding that OpenAI used freely available material, the interim status of the order, and the government’s proposed lawful-access test all make a broader conclusion premature.

    For now, companies do not need to predict exactly where Indian law will settle before improving their records. Source and rights information is far easier to preserve while a dataset is being assembled than after a dispute begins.

    RESEARCH CONTEXT

    This article draws on responses from Ankush Sabharwal of CoRover and lawyers Tanisha Khanna and Dinesh Jotwani. Legal findings were checked against the Delhi High Court’s July 24, 2026 interim judgment, the June 2025 and July 2026 orders in Bartz v. Anthropic, DPIIT’s December 2025 working paper, the European Commission’s training-content requirements, and recent US court filings. Company statements have not been independently audited.

    Read next: The Transformation Paradox — Why Organizational Readiness, Not Technology, Determines Whether Strategy Survives Disruption

    Topics

    More Like This

    You must to post a comment.

    First time here? : Comment on articles and get access to many more articles.