The training of large language models (LLMs) proceeds by ingesting vast corpora of text, a substantial portion of which is protected by copyright. This article asks a narrow but foundational question of Indian law: does that act of ingestion amount to "reproduction" within the meaning of Section 14 of the Copyright Act, 1957? Adopting a doctrinal and comparative method, the study disaggregates the training pipeline into six technically distinct operations and tests each against the statutory language of Section 14(a)(i), the definition of "infringing copy" in Section 2(m), and the storage-oriented amendments introduced by the Copyright (Amendment) Act, 2012. It argues that at least three of those operations — corpus acquisition, corpus curation and durable retention — constitute prima facie reproduction, that a fourth (tokenised batching) is defensible as transient and incidental, and that model weights are ordinarily not copies save in the narrow case of demonstrable memorisation. The article then evaluates the Delhi High Court's interim ruling in ANI Media Pvt. Ltd. v. Open AI OpCo LLC (2026), which held that such storage prima facie falls within the "private or personal use, including research" limb of Section 52(1)(a). It contends that this characterisation, while pragmatically attractive, strains the statutory text, conflates commercial secrecy with private use, and imports an open-ended proportionality analysis into a closed-list exception regime. The study concludes that the reproduction question in India cannot responsibly be resolved through interpretive elasticity and proposes a calibrated statutory text and data mining exception coupled with a collective licensing mechanism..