Right to be forgotten: how the delhi high courts decision fails against large language models and undefined public figures.
written by Sankalp Srivastava & Adithya Narayana Vojja both are 2nd year BA LL.B. (Hons.) student at Hidayatullah National Law UniversityAbstract
This article examines the superficial application of the Right to be forgotten by the Delhi High Court (DHC) in a batch of writ petitions, to protect the digital privacy of petitioners. Its sole reliance on de-indexing of URLs overlooks the ‘memorization’ within Large Language Models that encode data into model parameters for an indefinite period. Furthermore, the ambiguous exclusion of ‘public figures’, in the absence of a standard definition combined with inconsistent judicial application creates operational asymmetry. This piece advocates for a uniform regulatory mechanism that is capable of resolving the problems of doctrinal volatility and algorithmic data permanence.
Introduction
A settled legal principle that the law must be stable and yet it cannot standstill, aptly underscores the need for robust safeguards of informational privacy in the Digital Age. Against this backdrop, the concept of Right to be Forgotten was formulated that allowed individuals to request removal of their personal data from digital platforms when it is outdated, irrelevant, or harmful to their privacy. However, because of the absence of appropriate legal mandates in India, the question of enforcement of these rights remains unaddressed.
Recently, the DHC answered this question in a batch of petitions led by Laksh Vir Singh Yadav v. Union of India & Ors. By setting guidelines for enforcement of right to be forgotten including delinking/deindexing and masking of personal information from outdated judgements. These guidelines allowed petitioners to seek removal of URLs, restriction of name-based searchability, and masking of names from publicly available judicial records.
The guidelines dependency on deindexing, which is useful for static databases, is redundant in the case of Large Language Models as these models can retain personal data used for their training and reproduce it despite de-indexing/delinking of URLs. Effective implementation of these guidelines requires ‘machine unlearning’ which is costly and volatile. Moreover, the lack of a standard definition of ‘Public Figures’ combined with an indeterminate scope, causes haphazard exclusion of legitimate litigants. This vulnerability is further compounded by the inconsistent application of the ‘Public Role’ test accorded to similar petitions, which renders the guidelines unpredictable and inoperative in practice.
The Algorithm Does Not Forget: Re-evaluating the Right to be Forgotten in the Era of Generative AI
The Delhi High Court’s decision rests on the premise that de-indexing by search engines sufficiently addresses the privacy concerns. In doing so it ignored the artificial intelligence training systems, namely ‘Large Language Models’ (LLM). These LLMs are trained on vast amounts of text data designed to allow them to learn the structures and patterns of natural language and generate human-like text. De-indexing of URLs works for standard databases due to the static nature, but the same cannot be said of LLM’s as they synthesize answers based on learned patterns. Removal of external hyperlinks does not address the problem of internally absorbed data by these models, as research shows that precise training data can be extracted through targeted queries. Judicial data and records have been used to train these models, and this leads to ‘memorization’ of data, including full names and other personal identifiers whose verbatim can be extracted even after original public URLs have been disabled. To enforce the right to be forgotten effectively against these systems requires ‘machine unlearning’, which is unstable and comes with its own set of consequences. Computer science research reveals that erased data can resurface during subsequent downstream fine-tuning, even when there is no relation between the subsequent task and unlearning objective. This instability leads to enormous costs and complex programmes which is not feasible to the companies. The European Data Protection Board’s (EDPB) guidelines, explicitly states that developers cannot invoke these reasons to circumvent the rights of data subjects provided under Article 17 of GDPR. In the present case the Courts orders create a regulatory loophole as AI developers can permanently embed sensitive judicial data into model weights, using technical complexity as a cloak. Therefore, mere shielding of individuals from traditional name-based searches through search engines is not sufficient, as it leaves their privacy permanently vulnerable to extraction
A Test Without a Threshold: Deciphering the Ambiguity of the 'Public Figure' Exception in Indian Privacy Law.
On paper, the set of guidelines provided by the DHC appears to uphold the Right of Privacy by providing a remedial mechanism. However, the exclusion of ‘Public Figures’ from seeking remedies where ‘the allegations remain relevant to public discourse’, highlights a fundamental ambiguity as the judgment fails to specify the scope of a ‘Public Figure’.
In the absence of any statutory definition of ‘Public Figure’, the metric for determining the degree of popularity varies on a case-to-case basis. In Titan Industries Ltd. v. Ramkumar Jewellers , the Delhi High Court defined a celebrity as “a famous or a well-known person who "many" people talk about or know about”. Similarly, in Arun Jaitley v. Network Solutions Private Limited, the High Court observed that the popularity or fame of an individual will be no different on the internet than the reality. However, both the judgments fall short of setting a standard threshold for identifying a ‘Public Figure’.
This ambiguity is further exacerbated by the inconsistent application of the test laid down by the judgement, namely ‘if a person has voluntarily entered public life, their conduct in their public role is a legitimate subject of public scrutiny’. For instance, in the present batch of petitions, a petitioner who was a globally recognised figure in the fight against HIV/AIDS was granted relief on allegations relating to the HIV treatment administered by him, on grounds of his discharge and absence of any continuing public interest. Conversely, in a similar petition by a famous television artist, seeking removal of media depicting drunken behaviour, was rejected because of his public status and his alleged conduct being in the public domain.
The court held that Right to be forgotten is for protecting private individuals against information whose legal or social foundation has been extinguished and that mere passage of time does not extinguish the public interest in the conduct of a public figure. However, this reasoning reverts to the same issue of a standard threshold for determining a public figure. Such conflicting application of the public role standard reveals a major inconsistency which undermines the efficacy of these guidelines.
Towards Enforceable Forgetting: A Regulatory Roadmap for Algorithmic Privacy in India
It is clear that trying to alter permanent model weights is not feasible, therefore to bridge the gap between technical limits and privacy, the state should maintain a centralized registry of all de-indexed cases. The developers must then be legally bound to sync their models with this registry. Companies must ensure that targeted queries go under an external compliance layer, so that when a user attempts to extract protected data synced with the registry, this automated mechanism would block or redact the personal identifiers before giving the final response. This framework prioritizes regulation of limited visible output over the unstable machine unlearning. The government must designate these developers as Significant Data Fiduciaries under the Digital Personal Data Protection Act, 2023 (DPDPA). Since these developers and companies process data for providing AI services to domestic Data Principals, they come under India’s extraterritorial jurisdiction under Section 3(b) of the DPDPA, which legally binds them to the centralized registry framework. By doing so the liability is defined based on the control they have over data processing, and any non-compliance will lead to forfeiture of the traditional safe-harbour protection of intermediaries under Section 79 of the Information Technology Act, 2000 and also trigger penalty Section 33 and blocking order under Section 16 of DPDPA. The above framework combined with strict algorithmic data scraping mechanism in the National Judicial Data Grid, will create a robust system of accountability along with enforceability.
For the issue of ambiguity in ‘Public Figures’, Pro tem clarification by the Supreme court demarcating the scope of the term ‘Public Figures’ coupled with setting a standard procedure for the application of the “Public Role Test” would significantly reduce the ambiguity and operate as a precedent for lower courts.
Moreover, the government should be advised to formulate a comprehensive framework codifying the Right to be Forgotten within the Digital Personal Data Protection Act, 2023, comprising clear benchmarks for identifying public figures, public role of individuals and establishing a definitive criterion for delinking/deindexing and masking. In addition to this, public figures should be categorized according to the nature of their publicity, namely, politicians or celebrities with widespread fame should be categorized differently from people gaining fame accidentally or from specific controversies.
Lastly, a uniform test should be formalized for determining the “Public Figure” status with considerations such as Nature of public exposure, Voluntariness, Nexus between the information and public role and relevancy of public role.
Conclusion
The recognition of the Right to be Forgotten by the DHC is a significant step for Informational Privacy in the Digital Age, however the lack of measures for effective safeguards and structural inconsistencies renders it inadequate. Their reliance on de-indexing of URLs for protection of personal information ignores the ability of these models to retrieve confidential data despite it. Simultaneously, the ambiguity around Public Figure due to its lack of a standard definition and inconsistent application within the judgement, also demonstrates regulatory inconsistency. Addressing these deficiencies requires a centralized registry mandating compliance for screening of queries, strict algorithmic data scraping mechanism in the National Judicial Data Grid and integration of Right to be Forgotten in Digital Personal Data Protection Act, 2023. Conclusively, the validity of Right to be Forgotten will depend not on its recognition, but on its enforcement.
Subscribe now to keep reading and get access to the full archive.
