In November 2024, Asian News International (ANI) — one of India's largest news wire services — filed suit against OpenAI in the Delhi High Court, registered as CS(COMM) 1028/2024. ANI's claim is direct: ChatGPT was trained on ANI's copyrighted news reports without a license, the model reproduces ANI's reporting closely enough to substitute for reading the original, and in some instances it attributes fabricated quotes to ANI as if the agency had said things it never said. On 24 July 2026, Justice Amit Bansal of the Delhi High Court dismissed ANI's application for an interim injunction, holding — prima facie, not as a final judgment — that OpenAI's storage of ANI's content to train ChatGPT falls within Section 52(1)(a)(i)'s "private use, including research," and that ANI had failed to show that ChatGPT's outputs actually memorized or regurgitated its reporting. The suit itself is still alive; only the injunction application has been decided. India has run a fully digitized, high-volume, English-and-vernacular news and publishing industry for two decades, and its 1957 Copyright Act — amended piecemeal since then but never rewritten for machine learning — still has no provision that says in so many words whether feeding that corpus into a transformer's training loop is lawful; what changed in mid-2026 is that a court finally construed the existing text against that exact question, on an interim basis. You are about to learn how that reading was reached, how other courts have ruled on the same underlying question, and where India's statute is actually more permissive than most people assume, in one specific and surprising place: who owns what a machine writes.
What copyright protects, and what it does not
Section 13 of the Copyright Act, 1957 grants copyright in original literary, dramatic, musical and artistic works, among other categories. Section 14 lists the bundle of exclusive rights that copyright confers — to reproduce the work, to issue copies, to communicate it to the public, to make derivative works. Critically, none of this protects ideas, facts, or information — it protects only the particular expression of them. This is the idea-expression dichotomy, and in India it was settled doctrine well before AI existed, in R.G. Anand v. Delux Films (AIR 1978 SC 1613), where the Supreme Court held that two works can share the same theme or plot without one infringing the other; infringement requires that a substantial part of the original's expression be reproduced, judged by whether an ordinary observer viewing both works would get the unmistakable impression that one is a copy of the other. That "ordinary observer" test is qualitative, not mechanical — there is no percentage threshold in the statute — but it is the test an Indian court would actually apply to an AI output accused of infringement, and you will use a quantitative proxy for it below.
The second load-bearing concept is originality. Not everything written down is protected — the work must show some minimum creative contribution. India's Supreme Court addressed this directly in Eastern Book Company v. D.B. Modak ((2008) 1 SCC 1), a case about whether a publisher's edited, paragraph-numbered version of Supreme Court judgments could be copyrighted. The Court rejected the low UK "sweat of the brow" standard (protection for any labor at all) and also declined to adopt the US's stricter Feist "modicum of creativity" test wholesale, settling instead on a middle standard borrowed from Canada's CCH Canadian Ltd. v. Law Society of Upper Canada (2004): the work must reflect an exercise of the author's own "skill and judgment," not merely mechanical or trivial effort. This matters for AI on both sides of the pipeline — it sets the bar a scraped news article had to clear to be protected as a training input, and it is the same bar a prompt-plus-output pair has to clear if a human wants to claim authorship of what the model produced.
Is training on copyrighted text even legal? The fair use / fair dealing split
The training-data question turns on whether ingesting copyrighted works to compute statistical parameters counts as an infringing "reproduction" under Section 14, and if so, whether an exception applies. The United States and India answer this through structurally different tests.
US copyright law uses fair use under 17 U.S.C. §107 — an open, four-factor balancing test (purpose and character of the use, including whether it is "transformative"; nature of the copyrighted work; amount used; effect on the market for the original). Because the factors are balanced rather than enumerated, courts have room to find that large-scale statistical training is a new, transformative purpose distinct from reading the work — provided it does not substitute for the original in its market.
India's Section 52(1)(a) uses fair dealing, and fair dealing in India is a closed list, not a balancing test: dealing with a literary work is permitted only for (i) private use, including research, (ii) criticism or review of that work or another work, or (iii) reporting current events. There is no catch-all "transformative use" category. On a narrow reading, training a foundation model on millions of scraped articles maps onto none of the three: it is not private research in the ordinary sense (the resulting model is deployed commercially to millions of users), it is not a criticism or review of any specific ingested article, and it is not reporting current events. That narrow reading is not the only one on record, though. As you will see below, a Delhi High Court judge has already read "private use, including research" broadly enough to cover AI training, at least on an interim basis — so large-scale AI training in India is no longer simply unaddressed; it is contested, with one live judicial reading in OpenAI's favor and no final answer yet.
Other jurisdictions closed this gap deliberately, by statute. The EU's Digital Single Market Directive (2019/790), Article 4, creates an explicit text-and-data-mining exception that rightsholders may opt out of by reserving their rights (e.g., via a machine-readable robots.txt-style signal). Japan went further: Article 30-4 of its Copyright Act permits use of works for machine learning with no opt-out at all, provided the use does not "unreasonably prejudice" the rightsholder's interests, on the theory that a model consuming text for statistical pattern extraction is not consuming the work's expressive value the way a human reader does. India has amended the 1957 Act several times since 1994 but has added no equivalent TDM-specific provision, so an Indian AI company today cannot point to a statute the way a Japanese one can, nor invoke a US-style open transformative-use argument, because Section 52(1)(a)'s list is exhaustive by judicial interpretation.
How courts have actually ruled
On 24 July 2026, Delhi High Court Justice Amit Bansal decided the first real round of ANI v. OpenAI. He dismissed ANI's application for an interim injunction, holding — prima facie, on the standard used to decide whether an injunction is warranted, not as a final judgment after a full trial — that OpenAI's storage of ANI's copyrighted reports to train ChatGPT falls within Section 52(1)(a)(i)'s "private use, including research," and that ANI had failed to show that ChatGPT's outputs actually memorized or regurgitated its reporting rather than merely discussing the same news. He also rejected OpenAI's objection that the Delhi High Court lacked territorial jurisdiction merely because OpenAI's servers sit in the US. The underlying suit is still alive — the order decides only whether an injunction should issue while the case proceeds, not whether OpenAI ultimately wins — but it is the first time an Indian court has actually construed Section 52(1)(a) against a training-data fact pattern, on both sides of the pipeline: the input side (does storing text to train a model count as "research"?) and the output side (does the model's output reproduce the work?).
That reading cuts the opposite way from the most instructive foreign precedent, Thomson Reuters Enterprise Centre GmbH v. Ross Intelligence Inc. (D. Del., decided February 2025; on appeal to the Third Circuit as of mid-2026), where Judge Stephanos Bibas ruled on summary judgment against Ross Intelligence's fair use defense. Ross had trained a competing legal-search AI on Westlaw's copyrighted editorial headnotes — not on the underlying case law itself, which is not copyrightable, but on West's original summarizing and categorization of it. The court held the use was not transformative in the relevant sense, because Ross's output served the same market function as Westlaw: helping lawyers search case law. That directness — same purpose, same customers, same market — defeated fair use on both the first factor (purpose) and the fourth (market harm). Both cases turn on the same underlying question — does the AI's output substitute for the copyrighted source in its own market — but reach opposite provisional results: Ross lost because its product was a direct substitute for Westlaw, while OpenAI provisionally held its ground because ANI could not show that ChatGPT's answers reproduced its actual reporting rather than just the underlying facts.
Two other threads were still active as of mid-2026 and worth knowing by name rather than by outcome, since neither had reached final judgment: The New York Times v. OpenAI and Microsoft (S.D.N.Y., filed December 2023, now in summary-judgment briefing), alleging both training-data infringement and verbatim output regurgitation of paywalled Times articles; and Getty Images v. Stability AI, litigated in parallel in the UK and Delaware. In the UK action, Getty abandoned its primary infringement claims after the evidence closed at trial, because the training itself had taken place outside UK jurisdiction and copyright law is not extraterritorial. The claim that did reach judgment was a secondary-infringement theory: that Stable Diffusion itself was an "infringing article" imported into the UK. The High Court's November 2025 judgment rejected that claim too, holding that an infringing copy must at some point have stored or contained a copy of the work, and Stable Diffusion's weights do not. A separate, much narrower trademark claim succeeded — only for early model versions whose outputs reproduced visible Getty watermarks — but that is a Trade Marks Act 1994 finding, not a copyright one, and not a ruling on the core "is training infringement" question. The lesson generalizes: a company can locate its training compute in a jurisdiction with a favorable rule and largely sidestep another country's copyright law — which is exactly why the Delhi High Court's willingness to assert jurisdiction over OpenAI, despite its servers sitting outside India, matters as much as its fair-dealing holding.
Worked example: measuring "substantial similarity" quantitatively
Courts apply the ordinary-observer test qualitatively, but researchers and legal teams building a memorization or infringement screen — to flag candidate outputs before a human reviews them — need something computable. A standard proxy is n-gram overlap between a model's output and a candidate source text: the fraction of contiguous n-word sequences shared between the two, measured as a Jaccard similarity. This is not the legal test itself, but it is exactly the kind of first-pass evidence an infringement audit or a training-data deduplication pipeline (the technique behind Lee et al., "Deduplicating Training Data Makes Language Models Better," ACL 2022) actually computes.
def ngrams(text, n):
tokens = text.lower().split()
return set(tuple(tokens[i:i+n]) for i in range(len(tokens) - n + 1))
text_a = "the quick brown fox jumps over the lazy dog near the river bank"
text_b = "a quick brown fox jumps over the lazy dog beside the river"
A = ngrams(text_a, 5)
B = ngrams(text_b, 5)
overlap = A & B
union = A | B
jaccard = len(overlap) / len(union)
print(len(A), len(B), len(overlap), len(union), round(jaccard, 4))
Trace it by hand. text_a tokenizes to 13 words, so it yields 13 − 5 + 1 = 9 five-grams. text_b tokenizes to 12 words, yielding 8 five-grams. Comparing the two sets, four five-grams are identical in both: (quick,brown,fox,jumps,over), (brown,fox,jumps,over,the), (fox,jumps,over,the,lazy), and (jumps,over,the,lazy,dog) — the shared middle clause. The opening ("the quick" vs. "a quick") and closing ("near the river" vs. "beside the river") differ by one word each, which is enough to break every 5-gram that spans those positions. So overlap = 4, union = 9 + 8 − 4 = 13, and jaccard = 4/13 ≈ 0.3077. The printed line is exactly 9 8 4 13 0.3077.
A 30.8% five-gram overlap on a 12-13 word pair is a meaningful signal precisely because five consecutive words matching exactly is statistically improbable by coincidence in natural language — which is why Carlini et al. ("Quantifying Memorization Across Neural Language Models," ICLR 2023) define verbatim memorization using long exact-match spans rather than single-word or bigram overlap: short spans generate false positives from ordinary shared phrasing, long spans do not. That distinction is not academic; it is the difference between "this output happens to also use the phrase 'over the lazy dog'" and "this output reproduces the source's actual sentence."
Common misconception: "if it's paraphrased, it can't infringe"
A student's natural instinct is that copyright infringement requires word-for-word copying, so a model that rewrites a copyrighted paragraph in entirely different vocabulary must be safe. This is wrong, and it is wrong for the same reason non-literal software copying can infringe: copyright protects the expression, and expression includes structure, selection, and sequence, not only the literal string of words. R.G. Anand's ordinary-observer test explicitly asks whether the overall impression is one of copying, which a close paraphrase that preserves the source's distinctive narrative arrangement, argument order, and stylistic choices can trigger just as much as a verbatim quote — while conversely, an AI summary that extracts only the underlying facts (who did what, when) in a genuinely new structure infringes nothing at all, because facts are never protected regardless of phrasing. The n-gram screen above catches only the crude case (exact-string leakage); a real infringement analysis has to ask the harder, structural question the ordinary-observer test is actually built for, which no n-gram counter answers.
India's peculiar rule on AI authorship
Here the Indian statute is unusually forward-looking, almost by accident. Section 2(d)(vi) of the Copyright Act, added by the 1994 amendment (effective 1995) to cover works like database compilations and spreadsheet-generated tables, defines the "author" of a computer-generated literary, dramatic, musical or artistic work as "the person who causes the work to be created." India and the UK (Copyright, Designs and Patents Act 1988, s.9(3), which uses near-identical language) are among the very few jurisdictions that grant copyright authorship for machine output to a human at all.
Contrast this with the United States, where the D.C. Circuit affirmed in Thaler v. Perlmutter (2025, affirming the D.D.C.'s 2023 ruling) that a work with no human author is not copyrightable under the Copyright Act's use of "author" — a term the courts trace back to Burrow-Giles Lithographic Co. v. Sarony, 111 U.S. 53 (1884), which first held that a photograph could be copyrighted because the photographer made human creative choices (posing, lighting, timing), not because the camera did. The US Copyright Office's 2023 registration guidance formalizes this: AI-generated material is registrable only to the extent a human made copyrightable creative choices in selecting or arranging it, and purely AI-generated portions must be disclaimed.
India's Section 2(d)(vi) does not require that kind of human creative selection at all — it asks only who caused the work to be created, which on its face could cover someone who typed a prompt. But no Indian court has yet tested this against a generative AI output. The provision was drafted for deterministic, rules-based computer generation, not for a stochastic transformer sampling from a learned probability distribution, and it is genuinely unresolved whether "causing" a work by writing a three-word prompt clears even the low "skill and judgment" bar Eastern Book Company set for originality. India's law is more permissive in principle than the US's flat denial of AI authorship, but exactly how permissive is an open question, not a settled one — and until a court answers it, an Indian company relying on Section 2(d)(vi) to claim ownership of AI output is relying on an untested reading of a thirty-year-old provision.
Why this is a live economic question, not just a legal one
Major AI labs have already priced this risk rather than waiting for courts to resolve it: OpenAI signed content-licensing deals with the Associated Press (July 2023) and Axel Springer (December 2023), and Google licensed Reddit's data feed in February 2024, each paying to convert a disputed training-data claim into a clean contractual right. No comparable licensing deal exists yet between a major foundation-model company and an Indian publisher, which is one reason ANI went to court instead of the negotiating table — there was no established market rate to negotiate toward. For an Indian AI startup training or fine-tuning on scraped domestic content, the exposure is concrete: Section 55 provides civil remedies (injunction, damages, accounts of profits) for infringement, while Section 63 makes copyright infringement a cognizable, non-bailable criminal offense in India, unlike the purely civil US regime — a distinction that raises the stakes of getting the Section 52(1)(a) analysis wrong well above what a US-trained legal team might assume.
Active recall
Attempt each question before reading the worked answers that follow.
- Section 52(1)(a) of the Copyright Act, 1957 lists three purposes for which fair dealing with a literary work is permitted. Name them, explain why a narrow reading finds that none of them cleanly covers large-scale AI training-corpus construction, and explain how the Delhi High Court read the list differently in ANI v. OpenAI.
- Under R.G. Anand v. Delux Films (1978), what test does an Indian court apply to decide substantial similarity, and how would you apply it to an AI-generated summary that reports the same facts as an ANI article but in entirely different wording?
- Why does Thomson Reuters v. Ross Intelligence (2025) matter especially for AI tools built on a competitor's curated professional dataset (legal, medical, journalistic), rather than for general-purpose chatbots?
- India's Section 2(d)(vi) assigns authorship of a computer-generated work to "the person who causes the work to be created." Does this mean an Indian company can safely copyright ChatGPT's output today? Contrast with the US position in Thaler v. Perlmutter.
- Name two jurisdictions with an explicit statutory text-and-data-mining exception for AI training, and explain why India still has no statutory equivalent — and why that is no longer the same as saying Indian law is silent on the question.
- Rerun the worked n-gram example using 3-grams instead of 5-grams on the same
text_aandtext_b. Does the Jaccard similarity go up or down, and what does that imply about choosing n when using n-gram overlap as evidence of memorization?
Worked answers
- The three purposes are: (i) private use, including research; (ii) criticism or review, of that work or another; (iii) reporting current events. The textualist objection is that AI training fits none cleanly — the resulting model is deployed commercially to the public, so it is not "private" research in the ordinary sense; the training process does not criticize or review any specific ingested article; and building a general-purpose model is not "reporting" a current event. But in ANI Media v. OpenAI, Justice Amit Bansal read the first limb more broadly: on 24 July 2026 he held, prima facie, that storing copyrighted text to train ChatGPT is itself capable of being "research," because the storage is not disclosed to the public and ANI could not show the model's outputs reproduced its actual reporting. That holding decided only an interim injunction application, on a prima facie standard, in a suit that is still pending — a real, on-point judicial reading of the closed list, but not a final settlement of whether training belongs inside it.
- The court applies the "ordinary observer" test: after experiencing both works, would an ordinary person come away with the unmistakable impression that one copies the other? Because facts and ideas are never protected — only their expressive arrangement is — a summary that restates the same underlying facts (who, what, when) in genuinely different structure and wording would likely not infringe. It would infringe only if it also reproduced the original's distinctive phrasing, sequencing, or narrative structure, which pure fact-restatement does not.
- Ross trained a competing legal-search product on Westlaw's copyrighted editorial headnotes — original expression, not the non-copyrightable case law underneath. The court found this not transformative because Ross's output served the identical market function as Westlaw: legal search. Same purpose plus same customer base defeats fair use on both the purpose factor and the market-harm factor. Any AI trained on a competitor's curated professional dataset in law, medicine, or journalism sits in exactly this fact pattern — it is the same market-substitution theory ANI ran against OpenAI in Delhi, though unlike Ross, ANI could not point to any actual memorization or regurgitation in ChatGPT's outputs, which the Delhi High Court treated as decisive, at least at the interim stage.
- Not safely — the provision predates generative AI, was drafted for deterministic outputs like database compilations, and no Indian court has ruled whether typing a prompt constitutes "causing" a work to be created within Parliament's intended meaning, or whether it clears even the low "skill and judgment" originality bar set in Eastern Book Company. India's statute is more permissive in principle than the US, where Thaler v. Perlmutter flatly denies copyright to any work lacking a human author at all — but "more permissive in principle" is not the same as "settled," and an Indian firm relying on Section 2(d)(vi) today is relying on an untested reading.
- The EU's Digital Single Market Directive, Article 4, gives an opt-out TDM exception; Japan's Copyright Act, Article 30-4, gives a broader exception with no opt-out at all. India has not amended the 1957 Act to add any AI- or TDM-specific provision since 1994, and Section 52(1)(a)'s closed list still does not include text-and-data-mining as an enumerated purpose — so India has neither Japan's clear statutory yes nor a clear statutory no. But as of July 2026 it is no longer pure silence: at the interim-injunction stage in ANI v. OpenAI, the Delhi High Court read "research" under Section 52(1)(a)(i) to cover AI training-data storage, prima facie. That is a judicial reading, not a legislative one — provisional, confined to one case, and open to revision at trial or on appeal — so it narrows the uncertainty without closing it the way a TDM statute would.
- Recomputing with n=3:
text_a(13 tokens) yields 11 three-grams,text_b(12 tokens) yields 10. Six three-grams match exactly —(quick,brown,fox),(brown,fox,jumps),(fox,jumps,over),(jumps,over,the),(over,the,lazy),(the,lazy,dog)— since the shared middle clause now fully covers six overlapping 3-word windows instead of four 5-word ones. So overlap=6, union=11+10−6=15, Jaccard=6/15=0.40, i.e. 40%, versus 30.8% at n=5. Similarity increases as n shrinks, because shorter substrings are more likely to coincide by ordinary shared phrasing rather than actual copying. This means a low-n overlap score is weaker evidence of memorization — it is noisier and more prone to false positives — while a high-n exact match (Carlini et al. use spans of dozens of tokens) is much stronger evidence, which is exactly why memorization audits and legal evidentiary arguments favor long verbatim spans over short n-gram counts.
Think About It
Think about this: How would you explain ai and copyright law — training data, outputs, and indian ip to a friend who has never seen a computer? What real-world analogy would you use? Imagine you had to build a system using these concepts — what would be your first step? Try this: before moving on, write down three things you learned and one question you still have.