DPDP and AI: the Act never names it, and public data is not a free pass
By Abhijeet Singh · Primary sources verified by dpdprules.orgPublished · Last reviewed 15 min read
Does the DPDP Act apply to AI training data?
The short answer
The DPDP Act 2023 and the DPDP Rules 2025 never name artificial intelligence, machine learning or profiling. Each phrase appears 0 times in both. The framework still reaches AI, because section 2(x) defines processing as a wholly or partly automated operation on digital personal data and section 2(b) defines automated as any digital process capable of operating automatically, so training, indexing, storage and retrieval are processing whether or not anyone calls the system AI. The exclusion most often relied on is section 3(c)(ii), and it turns on who made the data public rather than on whether the data is public: it reaches personal data made publicly available by the Data Principal herself, or by a person under an obligation under a law in force in India to publish it, and a third party posting someone else's data is neither. The only provision that expressly names algorithmic software is Rule 13(3), which binds only a Significant Data Fiduciary, and as at 23 September 2026 none has been notified. There is no equivalent of a right against automated decision making anywhere in the Act. Section 3, section 9 and Rule 13 sit in the group due 18 months after publication, computed as 13 May 2027, which is interpretation until officially confirmed. Section 2 is already in force.

The phrase "artificial intelligence" appears 0 times in the Digital Personal Data Protection Act, 2023. It appears 0 times in the Digital Personal Data Protection Rules, 2025. So do "machine learning", "profiling", "inference" and "large language".
We measured those counts across the Gazette text of the Act, the India Code consolidated print, the Bill as passed by both Houses, the English Rules, the bilingual Rules extraction so that nothing could hide in the Hindi pages, and the corrigendum. The word "consent" was counted in the same pass as a control and returned 52 in the Act Gazette text, 53 in the India Code print, 53 in the Bill, 59 in the English Rules and 59 in the bilingual extraction, which is how we know the search was reading the documents rather than failing quietly. The corrigendum is a short list of typographical corrections and contains "consent" 0 times, so it carries no control of its own.
None of that means AI sits outside the framework. It means the framework reaches AI the same way it reaches cookies, through what is being processed rather than through the name of the thing doing the processing.
The Act reaches AI without naming it
Two definitions do the work, and both are already in force.
Section 2(b):
"“automated” means any digital process capable of operating automatically in response to instructions given or otherwise for the purpose of processing data"
Section 2(x):
"“processing” in relation to personal data, means a wholly or partly automated operation or set of operations performed on digital personal data, and includes operations such as collection, recording, organisation, structuring, storage, adaptation, retrieval, use, alignment or combination, indexing, sharing, disclosure by transmission, dissemination or otherwise making available, restriction, erasure or destruction"
Read the operative list against what building a model actually involves. Collection, recording, organisation, structuring, storage, indexing, combination, retrieval and use are each named. Assembling a training corpus is collection and combination. Holding it is storage. Building a vector index over it is indexing. Querying it at inference time is retrieval and use. On our reading every one of those steps is processing under section 2(x) when the corpus contains digital personal data, and nothing in the definition turns on the technique, the model architecture or whether anyone in the room calls the system AI.
Nothing in section 2(b) turns on how the instructions were produced either. A rules engine written by hand and a model whose weights were learned from data are both digital processes capable of operating automatically in response to instructions. The distinction that matters to engineers does not appear in the definition.
This is the same structure as the cookie question. The duty follows the personal data, not the storage mechanism and not the processing technique.
The one place an algorithm is named
There is exactly 1 provision in either instrument that expressly names algorithmic software. It is Rule 13(3):
"A Significant Data Fiduciary shall observe due diligence to verify that technical measures including algorithmic software adopted by it for hosting, display, uploading, modification, publishing, transmission, storage, updating or sharing of personal data processed by it are not likely to pose a risk to the rights of Data Principals."
Three limits are worth stating plainly, because guidance that quotes this rule often carries none of them.
It binds a Significant Data Fiduciary and nobody else. That status exists only where the Central Government notifies a Data Fiduciary or a class of them under section 10, and as at 23 September 2026 this site has located no such notification. Rule 13 is in the Rules' 18 month group under Rule 1(4), commencing on a computed 13 May 2027.
The duty is due diligence to verify that the measures are not likely to pose a risk to the rights of Data Principals. It is not a bias audit and not an explainability requirement, and neither of those appears anywhere in the framework. It is not the framework's accuracy standard either. That sits elsewhere: section 8(3) of the Act requires a Data Fiduciary to ensure completeness, accuracy and consistency where personal data is likely to be used to make a decision that affects the Data Principal or to be disclosed to another Data Fiduciary, and item (d) of the Second Schedule to the Rules carries a reasonable efforts version of the same standard for the processing that Schedule governs.
The full duty set attaching to that status lives on the Significant Data Fiduciary obligations guide, which is where the Rule 13 cycle, the reporting of significant observations and the localisation restriction are worked through. This article does not rebuild them.
Public data is not a free pass
This is where most of the published guidance goes wrong, and the error is worth isolating because it is an error about a single word.
Section 3(c)(ii) takes certain personal data outside the Act:
"personal data that is made or caused to be made publicly available by — ( A ) the Data Principal to whom such personal data relates; or ( B ) any other person who is under an obligation under any law for the time being in force in India to make such personal data publicly available."
Read what both limbs identify. Neither describes a state of the data. Both describe a publisher. The question the exclusion asks is not "is this data public" but "who made it public, and in which capacity".
That gives exactly 2 routes out:
- The Data Principal published her own personal data.
- Some other person published it while under an obligation, under a law in force in India, to do so.
A dataset scraped from the open web does not qualify merely by being on the open web. Where a third party posted someone else's personal data, that publisher is not the Data Principal, and unless an Indian law obliged them to publish it they are not within limb (B) either. The data is public and the exclusion does not reach it.
Limb (A) has a wider reach than a first reading suggests, because the words are "made or caused to be made" publicly available, which extends to a publication the Data Principal brought about without performing it herself. That reading, and how far it stretches over compiled and enriched records, is worked through on the startups applicability guide and is not rebuilt here.
Limb (B) carries a second constraint that is easy to read past. The obligation must arise under a law in force in India. A disclosure duty owed under a foreign statute does not satisfy the limb on the text, however genuine the obligation is in its own jurisdiction. We have found no Indian primary source addressing that point, so treat this as our reading of the words rather than as settled law.
Section 3(c)(ii) is also not the only route out, and an article about training data that stops there is incomplete. Section 17(2)(b) disapplies the Act to processing "necessary for research, archiving or statistical purposes if the personal data is not to be used to take any decision specific to a Data Principal and such processing is carried on in accordance with such standards as may be prescribed", and Rule 16 supplies those standards by pointing at the Second Schedule. Two things follow on the text. The exemption is conditioned on the data not being used to take a decision specific to a Data Principal, which is a real constraint on a model that will later be pointed at individuals. And the Second Schedule standards it imports include lawfulness, purpose limitation, data minimisation, retention limits, security safeguards and, at item (d), reasonable efforts to ensure completeness, accuracy and consistency. Whether a given training run is research within section 17(2)(b) is not resolved by any Indian primary source we have found, and we do not treat it as resolved here. Section 17 is in the same 18 month commencement group as section 3, and Rule 16 is in the Rules' 18 month group.
The 2 exclusions also operate differently, and the distinction matters to the downstream question. Section 3(c)(ii) is worded as taking a category of data outside the Act, because what section 3(c) does not apply to is "personal data that is made or caused to be made publicly available". Section 17(2)(b) is worded as disapplying the Act to a category of processing, because what section 17(2) does not apply to is "the processing of personal data" that is necessary for those purposes. On our reading it follows that an argument starting from the data has to be made about who published each record, while an argument starting from the purpose has to be made about the processing and has to satisfy the decision specific condition. How long that condition must keep being satisfied, and what follows if it stops being satisfied, is addressed in neither instrument and we have found no Indian primary source on it.
What the Illustration settles, and what it does not
The Act supplies its own worked example, immediately after the exclusion:
"Illustration. X, an individual, while blogging her views, has publicly made available her personal data on social media. In such case, the provisions of this Act shall not apply."
The Illustration is statutory text, not commentary, and it is the clearest thing in the framework on this question. It also has a narrow shape that is worth noticing: the person publishing and the person the data is about are the same person, and the consequence stated is about that data.
What it settles is limb (A) in its simplest form. X blogs her own views, and the Act does not apply to the personal data she thereby made available.
What it does not address is the question an AI training set actually raises, which is what a different person may then do with that data, at what scale, and for what purpose. The Illustration describes the publication. It says nothing about downstream reuse, and neither does the rest of section 3(c)(ii). On our reading the exclusion, once it bites, removes that data from the Act's application rather than licensing a particular use of it, which means the interesting arguments run to whether the exclusion was ever engaged for a given record rather than to what may be done afterwards. No Indian primary source has resolved this, and anyone who tells you it is settled in either direction is ahead of the material.
On our reading the practical consequence is that the test runs per record, not per dataset. A corpus of 10 million records is not inside or outside the Act as a block. Each record was published by someone, and limb (A) or limb (B) either reaches that publication or it does not. A dataset assembled from many sources will ordinarily contain a mixture, and on our reading the party relying on the exclusion is the party that has to show which records it reaches. Neither proposition is stated in the Act, and we have found no Indian primary source deciding either.
There is no counterpart to a right against automated decisions
The word "profiling" appears 0 times in the Act and 0 times in the Rules. There is no provision anywhere in the framework giving a Data Principal a right not to be subject to a decision based solely on automated processing, and no requirement to provide an explanation of automated logic.
The Data Principal rights chapter is Chapter III, and it confers 4: the right to access information about personal data in section 11, the right to correction and erasure of personal data in section 12, the right of grievance redressal in section 13 and the right to nominate in section 14. The Act confers 2 further rights outside that chapter, the right to withdraw consent in section 6(4) and the right of appeal to the Appellate Tribunal in section 29(1). None of the 6 is an automated decision right, and none of them turns on how a decision was reached.
A reader arriving from a GDPR background is usually looking for Article 22 at this point. That Article is the right place to look in the European instrument, and it should be read there rather than through this page. We do not restate GDPR requirements as claims to rely on. The claim being made here is about the Indian text, and it is a negative: the counterpart is absent.
The single place the framework addresses behavioural monitoring at all is section 9(3), which is confined to children:
"A Data Fiduciary shall not undertake tracking or behavioural monitoring of children or targeted advertising directed at children."
That is a prohibition rather than a consent question, which means no consent flow unlocks it, although Rule 12 and the Fourth Schedule exempt specified classes and purposes from it, and it is about children rather than about automated processing generally. The children analysis belongs to the children's data guide and is not developed here.
A research trap worth knowing
Search both instruments for the word "training" and you will get exactly 1 hit, and it is in the Rules:
"the Ministry of Personnel, Public Grievances and Pensions, Department of Personnel and Training"
It appears in the Sixth Schedule, printed under the words "See rule 21(2)", which sets the terms and conditions of appointment and service of officers and employees of the Board. It has nothing to do with training a model, and an argument built on that hit would be an argument about civil service staffing.
The same instrument holds a second near miss, and this one is closer to home. Search the stem "profil" and the Rules return exactly 1 hit, in the definition of "user account", which "includes any profiles, pages, handles, email address, mobile number and other similar presences". That is a noun about a social media presence. The word "profiling", the thing a reader searching that stem is actually looking for, still appears 0 times in either instrument.
This matters because these are the second and third traps of this shape we have documented. The word "succession" also appears exactly once in the Act, in section 18(2), where it gives the Board perpetual succession as a body corporate, and a search that finds it and builds an inheritance argument has cited the wrong provision entirely. A single hit on a promising keyword is a reason to read the surrounding clause, not a finding.
What this means in practice
None of the following is legal advice, and none of it is a requirement stated by the framework. It is what follows from the text above.
Record provenance at collection time, not later. Because the section 3(c)(ii) test runs on who published each record, a dataset whose provenance was never recorded cannot later be shown to fall inside the exclusion. Provenance is the evidence, and it is much cheaper to capture than to reconstruct.
Do not treat "publicly accessible" and "publicly available by the Data Principal" as the same category. They are different tests and only the second one is in the Act.
Check whether the Significant Data Fiduciary question is live for you. If it is, Rule 13(3) is the provision that reaches your algorithmic software, and the wider duty set comes with it.
Children are a separate question with a different answer. Section 9(3) is a prohibition, not a consent gate.
Watch the notification, not the commentary. The only things that will change this analysis are a commencement notification, a section 10 notification of Significant Data Fiduciaries, or an amendment. Trade body positions and consultation papers are neither.
When this starts to apply
Section 2 is already in force, which is why the definitions above can be stated in the present tense.
Section 3 is not. It sits in the group that paragraph (c) of commencement notification G.S.R. 843(E) places 18 months after publication, which this site computes as 13 May 2027 and labels interpretation until officially confirmed, because the notification states a period rather than a date and does not say how the period is counted. Section 9 is in the same paragraph, and Rule 1(4) of the Rules puts Rule 13 in the Rules' own 18 month group. Rule 1(3) is the one year group and holds only Rule 4.
There is a further wrinkle on the start date that this site tracks separately. The notification was published on 13 November 2025, but the eGazette file code on the same issue embeds 14112025, and whether the period runs from the 13th or the 14th is the subject of its own article. Nothing in this article turns on the 1 day difference.
So the applicability analysis in this article, including the exclusion everyone is arguing about, describes a provision that does not yet bite. That is not a reason to defer the provenance work, because the corpus being assembled today is the corpus that will be in place on the morning section 3 commences, and provenance not captured at collection time cannot be recovered afterwards. It is a reason to be precise about tense, and to be sceptical of any guidance telling you that you are in breach today.
What Rule 13 adds for a Significant Data Fiduciary →The section 3 applicability test →Why the duty follows the data, not the technology →Section 3, official text with sources →
Artificial intelligenceTraining dataSection 3Myth correction