Executive Summary
Posted to arXiv on September 23, a study from New York University asks why privacy keeps sliding to the back of the queue in the software that schools and universities run. The researchers interviewed 12 EdTech practitioners and scored the privacy policies of 48 platforms in active use, dividing each policy into five dimensions. The phrase the authors took for their title summarizes the sequence: we'll fix it later. This article follows that deferral into the public documents and asks which dimension it stops at.
The emptiest dimension was AI. Sixteen platforms out of 48, or 33%, run AI features a user can see on screen while offering no meaningful account of what that AI does. The other end of the rubric looks different. Seventy-nine percent spelled out the specific types of data they collect. The researchers also drew the boundary of the audit themselves. It covers only documents reachable in one or two clicks from a homepage, so a zero marks the absence of an accessible disclosure and not proof that nothing exists inside the company.
Sections 1 through 3 follow what the paper puts on the page. Section 4, which sets the findings beside the rules Korea now applies when a school picks learning-support software, is this article's own work and is not in the paper.
Key Figures
Source: arXiv 2609.28137 (September 23, 2026), Tables 3 and 4.
33%
Platforms with no AI disclosure
16 of 48. AI features were visible on screen and the primary policy documents explained none of them
73%
Middle score on accountability and breach
A general email address, vague breach language, and a claim of 'industry-standard' security with no detail
79% and 23%
Specific about collection, specific about AI
Measured against the same two-point bar, 79% cleared it for collected data and 23% cleared it for AI use
6.29
Average total across the 48 platforms
Out of 10, standard deviation 1.88. Each of the five dimensions takes 0 to 2 points and the five are summed
Deferring Looks Like the Right Call Every Time
That EdTech puts privacy off is not by itself a new finding. Earlier work has already recorded vague policies, excessive collection, uneven procurement, and organizational underinvestment. Separate strands have shown software teams treating privacy as a late-stage concern, developers of child-directed apps treating compliance as a threshold to clear rather than a foundation, and security underinvestment behaving as a rational choice when the incentives point the wrong way. What this paper adds is the part that is harder to see. It asks how those decisions actually get made inside an organization, then holds the answers up against the public documents.
The researchers contacted roughly 100 people, drew interest from 15, and completed 30- to 45-minute video interviews with 12. Three more could not take part because of non-disclosure agreements and internal compliance restrictions. Participants came from startups and nonprofits in the United States and India, a university research lab, and public K-12 districts. The authors make no claim that these 12 represent the industry. They write that they do not claim analytical saturation, and that their confidence rests only on the same accounts converging across people in different roles and contexts.
The first account that converged was deferral. Nobody said privacy did not matter. What repeated instead was a sequence: once the pilot is running, once the core functionality is in place. A participant who had co-founded a startup describes security as a second base layer, something the team would start putting in place once the pilots were going. Nonprofits and schools faced a different kind of pressure and landed in the same spot. Privacy work had to compete with visible program costs, so it needed a business case. A product manager at an adult learning organization said the team's first question was whether the investment would translate into student outcomes, growth, or future funding.
The second was delegation. A software engineer at a startup said the team relied heavily on Amazon's cloud for data storage, and a developer at another put every piece of data, user records and the company's own codebase alike, into Azure. Basic controls such as encryption and two-factor authentication were common. Threat modeling, provenance tracking, and systematic privacy review were described instead as work that turns urgent only after compliance pressure or an incident. Responsibility moved downward as well. One participant described smaller vendors bypassing district procurement and going straight to the school to ask for the data, a layer schools often have neither the equipment nor the understanding to manage.
The third was feedback. Usability ratings and comments on learning outcomes arrive steadily, and comments about privacy almost never do. A participant at an education nonprofit in India described a sandbox phase in which a small group evaluates the platform before full rollout, and said direct feedback on security and privacy features is uncommon, with most of it arriving as usability ratings. A participant at a university research lab traced the silence of users to a belief that one's own data could not be that valuable. Privacy harms are probabilistic, arrive late, and resist attribution, so an absence of complaints reads inside an organization as evidence that nothing is urgent.
Put the three together and every one of these decisions can look locally rational at the moment it is made. Cloud defaults cost less than building privacy engineering in house, pointing at a policy document is easier than operationalizing it, and waiting for complaints is cheaper than auditing data flows in advance. A principal data scientist at a large nonprofit left the paper a sentence that sums up this structure of time.
"Someone, oftentimes me, has to step in and say: it's much better for us to take our vitamins now than to have to take a serious dose of painkillers later."
The paper prices the painkillers in its opening pages. Unacademy in India exposed more than 20 million user profiles in 2020, Illuminate Education exposed sensitive records affecting more than 800,000 New York City students, and the 2024 PowerSchool incident reached tens of millions of records. For the second of those the paper cites a November 2025 announcement from the New York Attorney General's office. Three states, New York, California, and Connecticut, secured a $5.1 million settlement three years after the 2022 breach. What the investigation pointed to was not a sophisticated attack but the absence of basic measures, among them a failure to monitor for suspicious activity. The cost of a privacy failure, the paper explains, arrives late, spreads thin, and is carried by students more often than by the organization.
The authors add one more caution against reading this picture as good intentions alone. The deferral that surfaced in the interviews came in two kinds. In schools, nonprofits, and smaller organizations, staffing, budget, and expertise were frequently just absent. In some commercial settings, though, extensive behavioral tracking and personalization were tied directly to the value of the product, which makes collecting less data the same act as making the product worse. The university lab participant who explained the silence also spoke about cost, saying that maintaining privacy comes at a cost and that a new organization might have to forgo some of its profit if it does not collect the data or personalize the tool. Deferring for want of capacity and not deferring because collection pays do not yield to the same remedy.
AI enters this picture as an amplifier rather than a new item on the list. Where ownership of privacy is unclear, feedback is thin, and incentives point elsewhere, the arrival of AI makes what happens downstream even harder to inspect. The example the authors give is concrete. A parent may understand that a platform stores grades without knowing that a teacher's spreadsheet upload can send student data to a third-party model, or that an automated recommendation can reshape a learning path. That opacity weakens a feedback loop that was already limited, because a user cannot tell what processing took place, which actor made the decision, or where to raise a concern.
Disclosure Is Thorough Until the Data Moves
To see whether the sequence the interviews described also survives in public documents, the researchers scored 48 platforms. The set breaks down as 18 in US K-12, 14 in US higher education and adult learning, 11 in India K-12 and exam preparation, 2 in India higher education, and 3 tools that lead with AI. Selection favored market prominence within each segment, drawn from publicly reported usage, app-store rankings, and institutional adoption data. Segment sizes follow market maturity and the availability of public policy documentation, not an attempt at proportional representation.
There are five dimensions: what data is collected, what goes to third parties, how children and consent are handled, how AI and automated decisions are used, and who is accountable when something goes wrong. Each dimension takes 0, 1, or 2 points, designed to separate a topic being mentioned from a description someone could act on. A 2 on the AI dimension requires naming AI and stating what it trains on or decides. The first author built the five dimensions from the interview themes and from prior research on EdTech disclosure.
Two coders scored independently, then discussed and revised the places where their readings diverged. Agreement was strong but uneven across dimensions. Cohen's kappa runs between 0.86 and 0.88 for collection, third parties, and children, while the AI dimension sits at 0.654 and accountability at 0.645. The mean across the five is 0.781. The dimensions where two trained judgments diverged most were the dimensions where the documents themselves are vaguest. The paper states that all disagreements were by one point. Total scores from the two raters correlated at 0.865.
Set the adjudicated averages of the 48 platforms side by side, one bar per dimension, and the slope shows up at once.
Not one platform scored zero on collection, and 79% took the full two points. The AI dimension runs the other way: 16 of 48, or 33%, scored zero, which means nothing in the primary policy documents meaningfully explained an AI feature a user can see on screen. On accountability the trouble sits in the middle score instead of the zero. Seventy-three percent landed there, and the typical shape is one general email address, vague breach-response language, and a claim of industry-standard security carrying no specifics. The 19% that scored two named a contact, a notification timeline, and a standard.
One more dimension repays a second look. Children and consent averages 1.05, which places it near the middle, but it reaches that middle by a route the others do not take. Thirty-five percent scored zero and 38% scored two. It carries the highest share of zeros among the five dimensions and the widest spread of scores, the opposite shape from accountability, where the average fell because the scores bunched together. Within this one dimension, some platforms set out both the law and their actual practice in a section of its own while others never address children at all. The numbers in the next section explain what produces that split.
The paper attaches an explanation to the slope. Describing collection is commonly required by law and procurement, and above all it is easy to put into a document. Writing a meaningful AI disclosure means working out and recording model and data flows; writing a meaningful accountability clause means naming a person, a deadline, and a standard. What is left is a state in which public expectations are satisfied while the machinery for governing downstream use goes unbuilt. The paper calls this the compliance-practice gap.
The scope of the scoring belongs in the picture as well. Only documentation reachable in one or two clicks from a platform homepage was in range. Pages requiring a login or buried deep in a support portal were excluded. A zero therefore means no accessible disclosure within that scope, and the authors write in advance that it is not evidence of an absent internal commitment. The limitations section says the same of the audit as a whole, that it is a snapshot of public documentation and not proof of internal behavior. That scope, though, is also the range that schools and parents can actually open. We have looked at how much an audit of public policies alone can catch before, writing about a study that found contradictions inside the policies of 123 companies.
The Gains Stop at the Dimension the Law Names
Split by segment, the distance is plain. Out of 10 points, US K-12 leads at 7.47, followed by US higher education and adult learning at 5.50 and India K-12 and exam preparation at 4.77. The difference was statistically significant (Kruskal-Wallis H=20.634, p<0.001). The order itself is no surprise: the segments under heavier external pressure also score higher.
So where was K-12 ahead? The dimension on which K-12 platforms scored significantly higher was children and consent, the dimension the Children's Online Privacy Protection Act and the Family Educational Rights and Privacy Act govern directly. And the advantage did not carry over to AI disclosure or accountability. On the dimensions where no equivalent mandate applies, K-12 looked like everyone else.
[Fact] Regulation improves the dimensions it specifically makes enforceable rather than producing generalized privacy maturity, in the paper's words. [Interpretation] Not because organizations are bad, but because that is how the queue gets built. A dimension with a name acquires an owner and a deadline, and a dimension without one moves to next quarter.
This result reads as heavier in EdTech than elsewhere for a reason. The habit of putting things off exists in software everywhere, but the market ordinarily corrects it to some degree. People leave when they dislike something, reputation takes a hit, and the signal travels back into the product. In a school that machinery does not run. A student does not pick the learning management system or the classroom app, and a family may not know which services are in use at all. Harms such as profiling, secondary use, or an altered learning path accumulate quietly and resist being traced to any one decision. Where exit, reputation, and ordinary market feedback are all unreliable correctives, the only forces left to push improvement are procurement requirements and regulation.
The authors rule out reading these numbers as one country outscoring the other. The India segment is small, and enforcement of India's Digital Personal Data Protection Act was still finding its feet during the study period. Segment sizes and platform mix are uneven as well. The score difference between the United States and India is therefore a pattern observed within this sample and not a measurement of a country's privacy maturity, as the paper puts it.
Korea's Required Items Never Ask About AI
From here on we are outside the paper. The study looked at US and Indian platforms, and Korea was not in scope. Korea, though, is in the middle of setting new criteria for the software its schools use, which makes this a good moment to hold the paper's five dimensions up against them.
Article 29-2 of the Elementary and Secondary Education Act, added in August 2025, took effect on March 1, 2026. For a principal to adopt learning-support software built on intelligent information technology as instructional material, the school must follow criteria set jointly by the Minister of Education and the Personal Information Protection Commission, and the choice must pass review by the school operating committee, the statutory body of teachers, parents, and local members that each school convenes. The guidance on the selection criteria that the Ministry of Education published at the end of December 2025 splits those criteria into five mandatory items and five optional ones. Software that misses even one mandatory item cannot be used.
Placed next to those five dimensions, it comes out like this.
| The paper's audit dimension | The matching item in Korea's mandatory criteria |
|---|---|
| What data is collected | Minimal processing — whether collection is kept to a minimum, plus a statement of purpose, fields collected, and retention period |
| What goes to third parties | A statement of third-party provision and processing delegation (where applicable) |
| Children and consent | Safeguards including consent from a legal guardian for children under 14 |
| AI and automated decisions | No matching item |
| Accountability and breach response | Contact information for the privacy officer (the item carries no notification timeline or standard) |
Two further mandatory items sit alongside these: a duty to apply safeguards, and a procedure through which a user can demand access, correction, deletion, or suspension of processing at any time. That last one is a user-rights item the audit rubric does not have, and a place where the Korean criteria are the ones out front. One dimension is empty. Nothing among the five mandatory items asks whether AI is in use, whether a student's inputs or learning history feed back into model training, or which decisions an automated recommendation stands in for. Given that the statute aims at software built on intelligent information technology, that blank is conspicuous.
In May 2026 the Personal Information Protection Commission and the Ministry of Education opened a joint advance inspection under Article 63-2 of the Personal Information Protection Act. It covers the seven most-used services among those registered on the ministry's edtech registry and the digital tools chosen by metropolitan and provincial education offices, and the items checked are consent for collection and use, destruction once the purpose has been met, procedures for collecting children's data, and safeguards such as vulnerability testing and the management of access rights and access logs. The two agencies stated their aim as a move from managing incidents afterward to preventing them beforehand. Here too the items gather around collection, consent, and safeguards.
Korea's system rests on the privacy policy as well. The Q&A in the ministry's guidance tells schools that when the registry board does not show whether a product meets the mandatory criteria, they should ask the company directly or check the privacy policy on the product's homepage. The company checklist attached to the guidance carries examples in its evidence column that read "Article X of the privacy policy." What the paper recommends to schools and procurement staff takes aim at exactly that spot: do not treat the existence of a policy document as sufficient, and ask for operational evidence. And this study, by scoring 48 platforms, showed that the dimensions those policy documents write about least are AI and accountability.
The absence of an AI item from the Korean criteria is not a failure in itself. The regime has only just finished its first semester, and the optional criteria are left open for schools to extend on their own. But if the pattern the paper observed in those two countries also holds here, a dimension nobody names will not fill itself in. Obtaining even a list of which AI sits inside a school's software can be hard, as we saw when New York City switched off AI in 38 learning programs and would not name them. The question the authors left for future work runs the same way. They propose tracking over time whether procurement requirements or AI-specific rules spill past the dimensions they directly govern. Korea's five mandatory items mean a case for testing that question has just come into existence.
Why Pebblous Is Watching This Study
The last section of the paper is a prescription rather than a diagnosis. And that prescription overlaps almost entirely with the list we reach for when we explain data governance.
What the authors propose to developers and firms is that privacy become a release criterion instead of a documentation exercise. Concretely, a stage gate: before a feature reaches students, the feature owner writes a one-page data-flow summary. What data is collected, which third parties receive it, whether it is retained or used for model training, what automated decisions occur, and who owns incident response. They also write that the moment an educator uploads a roster or a spreadsheet, the interface should make third-party transfer and model-training use visible right at that point. To schools and procurement staff they say to demand operational evidence instead of accepting that a policy document exists, and they add that the five audit dimensions serve as a lightweight comparison rubric.
A separate recommendation goes to teachers. A teacher is often the final gate a tool passes through on its way into a classroom, yet may lack the time, the legal expertise, and any window into data flows, so the authors propose training on concrete decisions instead of leaving the matter to individual judgment: recognizing the moment student information is being transferred, checking whether AI use is disclosed, and knowing which institutional channel owns approval. Literacy cannot substitute for organizational accountability, they write, and students and teachers should not be expected to compensate for opaque vendor practices.
There is no new technology on this list. Collected fields, third-party recipients, whether data gets reused, automated decisions, the person accountable. What we call data lineage and an AI bill of materials comes down to making those five lines update every time the product changes. The NYU researchers did not write this prescription with any knowledge of the product category Pebblous works in. This article carries the paper over because they arrived at the same items after interviewing practitioners and scoring documents.
The questions below move the paper's five dimensions into the work of a product team, and they are not a checklist printed in any document.
- Which of the data we collect today has a user never seen on a consent screen? Are behavioral logs and error records on that list?
- Could the owner of a feature put on one page where the data entering our model comes from and where it leaves for?
- Can we tell a user inside the product, at the place where it happens, that their data is being reused for model training?
- Is who does what in the first hour after an incident written down with names and deadlines? A general inquiry address does not count.
The paper closes like this. Closing the gap calls for institutional procurement standards, clearer internal ownership, and enforceable oversight, and privacy has to become a condition of deployment rather than a promise to revisit later. Behavioral logs, disability records, and academic histories pile up before a child is old enough to consent to any of it. There is that much less time to defer the condition.
Thank you for reading this far. Every figure and sentence this article carries can be checked by anyone against the paper itself. Who in your organization could write that one-page data flow today? If no name comes to mind, what this paper heard in its interviews is not somebody else's story.
References
Academic Paper
- 1.Nair, M. M., Greenstadt, R. (2026). ""We'll Fix It Later": Education, AI, and the Deferral of Privacy in EdTech." arXiv:2609.28137.
Reported Breaches
- 2.Bannister, A. (2020). Data breach at Indian learning platform Unacademy exposes millions of user accounts. The Daily Swig (PortSwigger).
- 3.James, L. (2025). Attorney General James and Multistate Coalition Secure $5.1 Million From Education Software Company for Failing To Protect Students' Data. Office of the New York State Attorney General.
- 4.PowerSchool. (2024). SIS Incident – Notice of United States Data Breach. PowerSchool Security Notice.
Korean Policy
- 5.Ministry of Education, Republic of Korea. (2025). Guidelines for Selecting Safe and Effective Learning-Support Software for Schools. Ministry of Education notice. (Korean)
- 6.Newspim. (2026). Ministry of Education and PIPC to Conduct Joint Privacy Inspection of EdTech Firms. (Korean)