Executive Summary
This article reads a preprint posted to arXiv on 13 September 2026, unreviewed, through one question: once a conversation ends, what does the product write down about you? The study took conversation histories that users in India, Nigeria, Brazil and Pakistan donated with consent, and set them beside the profile entries the chatbot had pulled out of those same conversations and stored separately. The audited system is the legacy memory feature as it stood in February 2026, and for nineteen entries out of every twenty, no request to save came first.
One summary that usually attaches itself to a study like this gets no support from the paper's own tables. The share taken by symptoms and stigmatised conditions is smaller on the profile layer, not larger, and the proportion graded high risk is practically identical across the two layers. The paper's claim about risk does not rest on a rate. It rests on density, several cues collecting on a single line, and on permanence, that line staying on. Evidence that the density really does single people out already exists over these same users. The same first author published it four months earlier, and this paper does not cite it.
Open the product documentation and the target of the comparison moves. The sentence saying details may be saved without the user asking is still on the help page today, and four months after the audit window the background extraction was promoted to a design goal in a company announcement. That leaves two gaps. One is proportion: the path written as an aside turned out to be the main road. The other is a promise about health that is no longer in the current text. When a derived line has nowhere to record when it was observed, how long it holds, what it rests on and who asked for it, no retention period and no correction request can ever reach that line. And the route for redoing this count from outside had closed before the paper appeared.
95.36%
Profile entries with no save request in front of them
Denominator 7,051 profile entries. It means the single message before the save carried no save-instruction wording
41.11%
Profile entries carrying health talk
2,899 of 7,051. Inside them sit 135 stigmatised conditions and 68 medical record numbers
3.62% · 3.63%
High-risk share, conversation layer and profile layer
Over 179,057 conversations and 7,051 profile entries. The two differ by 0.01 percentage points
F1 0.90
Gender recovered from the same 1,057 users by a general-purpose model
On conversations filtered so that no message stated who the user was. Measured by a sister paper four months earlier
The line that was not in the reply stayed in the profile
The paper opens on a short scene. A user tells the chatbot their age and occupation, then adds that lately they keep forgetting things. The reply that comes back holds not one sensitive inference. It is an ordinary answer that takes the symptom in stride, so on screen it looks as though nothing was left behind. In that same moment the product pulled those facts out and stored them. The user never said to save anything, and no notice of the save appeared on that screen. The caption the authors put under the figure points at all three things at once: that the user disclosed demographics and a symptom, that the visible reply contains no sensitive inference and so gives the impression that nothing persisted, and that the stored entry quietly holds those facts and carries them into later conversations.
This scene is not a real conversation. The researchers composed it from what they saw in their own dataset. Every number that follows does come from real records, and those numbers say exactly the third item in the scene. So the vocabulary first. This article uses profile entry for the one-line storage unit that ChatGPT's memory feature writes when it pulls a single fact out of a conversation. The paper's own words are memory entry and profile entry. It is a different layer from the conversation log, and almost every comparison runs between those two layers.
1.1This is not a breach
First, what is not at issue. No data escaped. There was no intruder, no misconfigured server, no sign that anyone read someone else's conversations. The product here worked as designed. It pulled facts out of conversations and stored them separately, and that storing happened almost independently of what the user asked for. The paper counted those saves, and this article reads the count. Because the subject is default behaviour rather than an incident, the numbers that follow read differently.
1.2What was counted, and how
The material is ChatGPT export files donated with consent by users in India, Nigeria, Brazil and Pakistan. The donated raw corpus held 202,590 conversations from 1,252 people, and after dropping the bottom decile by message count, 179,057 conversations from 1,057 people went into the analysis. Of those, 766 people had files containing entries written by the memory feature, and the entries numbered 7,051 in total. The authors take the remaining users to have had the feature switched off. Classification ran on Llama-3.3-70B, an open-weight model, and the authors state they avoided closed APIs because of how sensitive the material is. Whether a save had been requested was judged from the single message immediately before the entry was stored, using trigger lexicons in four languages and an edit-distance similarity threshold of 80%.
A limit is attached to the subject. Collection happened in February 2026, and what was audited is the legacy memory feature as of that date. The paper brackets the limit itself: the body states that the documentation its own section 4.3 compares against is the help page for the legacy memory feature, "not the current dreaming architecture." The epigraph and the opening paragraph of the introduction, meanwhile, quote the announcement of the new architecture. The old and the new are not conflated. Only the door is the new one. This distinction becomes necessary again in section 6.
1.3How this differs from earlier pieces here
Two months ago this blog covered the problem of deleting memory an agent writes for itself. That piece also observed agents writing on their own judgement rather than on instruction. Two things differ. The evidence there came from observing the memory that agents use in development work, and the evidence set out below is a count over the records of 1,057 real users of a consumer product. The empty box differs too. That piece asked whose is this; the box counted in the sections below is when is this from. The two questions sit on the same line and ask different things.
Two more pieces are worth reading alongside. The piece on the implementation stack that makes a chatbot remember earlier conversations looked at how the feature gets built, and the sections below count what the built feature actually wrote down. And it sits on a different layer from the piece on memorisation bias in medical AI, published here the same day. The medical-AI piece is about records used for training persisting inside model weights, and this one is about sentences pulled from conversations persisting as text inside a product. The sister paper introduced later makes the distinction in its own words: inference is not memorisation, and the defences against memorisation, training-data deduplication, differentially private training and training-data opt-out, do not touch the inference side. That is the account given in its section 6.2.
One conversation in five was about health
Of the 179,057 conversations, 38,165 carried health talk, a rate of 21.31%. That is the value the abstract and the body put forward as the headline. Add the same table up in a different direction, though, and the value splits. The gender rows and the country rows each sum to 37,828, and only the age rows sum to 38,165. Added to the 141,229 conversations with no health talk, 37,828 hits the total of 179,057 exactly, while 38,165 leaves 337 over. Which figure is the typo cannot be settled from inside the paper. This article uses the authors' 21.31% as published, and records alongside it that the split which reconciles is the 37,828 one.
Risk was graded on a five-level rubric. Over the 179,057 conversations, score 2 accounted for 14.03%, score 3 for 3.27%, and scores 4 and 5 together, the high-risk band, for 3.62% (6,532 conversations). The definitions carry over exactly. Score 4 is where the place of care shows, meaning the name of a specific hospital, clinic or clinician. Score 5 is a medical record number or patient ID, or a condition that carries stigma. The test is which hospital and which clinician. It has nothing to do with anything a satellite would fix. Medical record numbers surfaced in 632 conversations. Within the conversations that carried health talk, symptoms and physical states took 16.02%, care locations 8.57% and stigmatised conditions 8.13%. The sociodemographic-proxy and lifestyle categories are larger than these, but their shares are not written as numbers anywhere in the body, only in a figure, so we do not carry those values over.
2.1Why these four countries
The four countries were not picked at random. The paper gives doctor shortages, crowded public clinics and high out-of-pocket medical costs as background, without ever citing the figures. So we pulled them from World Bank Open Data ourselves. The last two columns of the table below are not in the paper. They are our own lookup. The latest non-missing year differs by country, so the reference year sits next to each value.
| Country | Participants | High-risk conversations | Physicians per 1,000 | Out-of-pocket health spending |
|---|---|---|---|---|
| India | 456 (43.1%) | 3.2% | 0.72 (2020) | 43.9% (2023) |
| Nigeria | 206 (19.5%) | 3.4% | 0.38 (2023) | 71.9% (2023) |
| Brazil | 205 (19.4%) | 5.5% | 2.36 (2023) | 26.2% (2023) |
| Pakistan | 190 (18.0%) | 3.2% | 1.16 (2021) | 52.9% (2023) |
Participant counts and high-risk conversation rates come from arXiv:2609.14697 Tables 1 and 3. Physician density and the out-of-pocket share were queried directly from the World Bank Open Data API (SH.MED.PHYS.ZS, SH.XPD.OOPC.CH.ZS) on 17 September 2026.
Nigeria has the fewest physicians and the largest share of care paid out of pocket. Those two numbers together explain why a chatbot ends up standing in for a first consultation. The table does not resolve cleanly. Brazil, with the highest physician density of the four, also had the highest rate of high-risk conversations. Access alone does not account for that. This reading is ours and is not a conclusion of the paper.
Any read of subgroup differences has to keep sample sizes in view. Section 5.1.3 of the paper devotes a subsection to risk patterns among older users, where the 65-and-over group is one user with 5 conversations and the 55-to-64 group is eight users with 720 conversations. For the same reason, the statement that the 55-to-64 profile entries contain zero high-risk items reflects there being only nine entries rather than any filtering by the system, and the authors say so themselves. Country comparisons carry a language condition. The trigger lexicon that decides whether a save was requested exists in four languages only, and languages widely spoken in Nigeria, Hausa among them, fall outside the classifier, as the authors state in the methods section.
2.2This is not the first use of this corpus
These conversations were not gathered for this paper. A pipeline for recruiting participants through Clickworker, taking consent and receiving donated export files came first, and this paper reanalysed a corpus built under the same ethics approval. This is at least the third analysis to use it. The original paper reporting the collection itself survives in the bibliography as a title with no venue and no URL, and the citation year is 2018, before ChatGPT existed. We could not find a published version either.
Last month this blog covered the problem of measuring chatbot use with an independent corpus, and wrote this about the coverage limits of such measurement: private deployments, on-device assistants and services widely used across the Global South mostly fall outside the frame. This paper steps into exactly that space. On the side of who confides health talk to a chatbot, the piece on mental-health conversations among US teenagers came first. That story looked at who is speaking. The question left over is where what was said gets written down, and in what form.
2.3A de-identified corpus where hospital names define high risk
The corpus design and the scoring rubric collide in one place. Right after collection, the corpus was de-identified by replacing names, emails and phone numbers through named-entity recognition and regular expressions. Yet the definition of the high-risk grade is hospital names, clinician names and medical record numbers. The categories chosen for removal and the categories chosen for counting overlap. So the 3.62% is either lower than the true figure, or it is the rate at which de-identification missed. This point is not in the paper. It is our reading.
The answer to that question already exists. In May 2026, four months earlier, the same first author published a different paper using the same donated corpus, the same ethics approval, the same 1,057 people and the same Llama-3.3-70B. Here is what that paper measured. After keeping only conversations in which not a single message stated who the user was, an unremarkable general-purpose model recovered age, gender and country at weighted F1 of 0.84, 0.90 and 0.88. A sentence from its section 6.1 reads as follows.
De-identification did not fail. The parts it never touches were enough on their own. The abstract and section 6.1 of that sister paper describe the cohort filter slightly differently from each other, so only the conservative form is used. It says conversations filtered so that no message stated who the user was, and it does not assert that this held despite de-identification. And this paper does not cite that sister paper. Searching the body and all 41 bibliography entries turns up neither the title nor the arXiv number. The intent is not ours to know. Two things can be written down: that the citation is absent, and our reading that including the result would have changed the conclusion of section 4.
Where the documentation took a promise back
Of the 7,051 profile entries, 327, or 4.64%, had a user save instruction immediately before them. The other 95.36% were made where nobody had asked. The authors report no significant difference by country, by gender or by age band. In the bar below, the orange portion is the entries the user brought about.
To avoid reading this number as a measurement of consent, the method has to be precise. The researchers looked at the single message immediately before the entry was stored, and at nothing else. They searched that message with regular expressions against a save-instruction lexicon built in four languages, English, Portuguese, Hindi and Urdu, and allowed up to 80% edit-distance similarity to absorb typos. The 80% threshold came from a human review of 200 cases: go lower and everyday questions like "do you remember that news?" get dragged in, go higher and human typos slip past, as the authors explain. So what the 95.36% says is that the one message before the save carried no save-instruction wording. Whether switching the feature on in settings amounts to blanket consent, and whether the user had already asked earlier in the conversation, cannot be separated by this method. And languages such as Hausa are not in the lexicon.
The rate itself is not this paper's first sighting either. Earlier work posted to arXiv in February 2026 and published at the Web Conference that April looked at 2,050 memory entries from 80 users and already reported that 96% were system-generated alone. This paper describes that study as a "similar finding" and as a "confirm." A separate study with a sample more than thirteen times smaller produced a neighbouring rate, and this paper's contribution lies not in the rate but in the health domain, the four countries, and the comparison across the conversation and profile layers.
3.1The paper's appendix contradicts its own figure caption
The paper sets this 95.36% against the product documentation. Inside Appendix B, however, two statements diverge. The appendix body conveys that the help page still says something like "remember that" has to be uttered for a save to happen, while the caption on the same appendix figure says that page states details may be saved without being asked. One of the two is wrong, and settling it means reading the source document ourselves. So we did. As of 17 September 2026 the help page reads as follows.
"Remember that" appears as an example, not a requirement, and the very next sentence says details may be saved without the user asking. The caption is right and the appendix body is wrong. This is also the one place where this article ends up more accurate than the paper. That verdict moves the target of section 3. The comparison of "the documentation said user-directed, the reality was automatic" does not hold up. That saves also happen automatically is written in the documentation. The February 2024 post that first announced the memory feature already said the product would either be told directly or pick things up on its own.
3.2So two gaps survive
The first is proportion. The documentation lists two paths side by side: being told directly, and picking things up on its own. A reader takes the first for the main road and the second for an aside. The measurement says the opposite. Of 7,051 entries, 95.36% came by the second path. The documentation did not state a falsehood. Its word order simply runs against the actual distribution. Something else diverges in that contrast too. The litigation over an order to stop deleting conversation logs, covered on this blog in July, dealt with company statements made in the course of a lawsuit. This time the counterpart is a sentence in product help that is public right now.
The second is the exception promised for sensitive categories. The February 2024 announcement had a paragraph on privacy, and inside it a sentence that named health specifically.
The verb is "steer away from," and ahead of it sits "we're taking steps to." The "trained not to proactively remember" form that circulates in search summaries is a paraphrase of this sentence rather than the sentence. Here is the measurement to set against that promise. Of the 7,051 profile entries, 2,899 (41.11%) carried health talk, and inside them sat 135 stigmatised conditions and 68 medical record numbers. The promise belongs to a February 2024 document and the measurement to February 2026 data, so setting the two numbers down together requires attaching the dates. And the current help page does not contain that promise. In its place sits a sentence of the opposite polarity: "Sensitive information may appear in memory if you share it with ChatGPT." The paper's appendix likewise notes that the company later removed some of the wording that had emphasised user control.
This comparison rests on the text as it stands on 17 September 2026. The version of the help page from February 2026, when the researchers collected their data, could not be checked, because the Internet Archive had suspended service at the time of our lookup. The documentation screenshots the paper carries in its appendix were themselves taken over from the earlier study mentioned above and date to January 2026.
3.3Four months later, that behaviour became a selling point
The audited data is from February 2026 and the paper appeared in September, and in between, on 4 June, the company announced a new memory architecture. One sentence in that announcement is the most valuable thing in this article.
The behaviour the paper called a gap had, four months later, been promoted to a design goal in a company announcement. Not relying on explicit requests is now a property worth advertising. The paper did not make this comparison. We read the two documents against each other. We also counted words. Search that announcement's body for sensitive, health, privacy, consent, opt out and export, and every one of them returns zero. Privacy matches three times, all of them navigation links at the foot of the page. The February 2024 announcement had a privacy paragraph. The 2026 announcement, which raised background extraction into a foundation, does not. This count is also ours.
On which layer is it true that the risk went up?
The paper's abstract says that memory synthesis running in the background increases re-identification risk. Yet take three rulers to the same paper's tables and the three point in different directions. The share of health talk goes up, the share scoring 2 or more on risk goes down, and the high-risk share barely moves. The table below gathers all three. Two of the three profile-layer values are ones we recomputed from counts the paper printed.
| Measure | Conversation logs (179,057) | Profile entries (7,051) | Direction |
|---|---|---|---|
| Share carrying health talk | 21.31% (38,165) | 41.11% (2,899) | up |
| Risk score 2 or higher | 21.13% (37,828) | 18.31% (1,291, computed) | down |
| High risk (4 and 5) | 3.62% (6,532) | 3.63% (256, computed) | practically equal |
| Symptoms and physical states | 16.02% | 4.31% (125) | smaller |
| Stigmatised conditions | 8.13% | 4.66% (135) | smaller |
| Care locations | 8.57% | 6.45% (187) | smaller |
Carried over from arXiv:2609.14697 §4.1, §4.2 and Table 5. The two marked values are ones we computed from the counts the paper printed: risk score 2 or higher is 7,051 minus the 5,760 graded no-risk, and high risk is the 160 at score 4 plus the 96 at score 5. The last three rows use health-talk items on each layer as the denominator, so their base differs from the first three rows.
Putting the first three rows into bars makes the directions easier to weigh against each other. The grey bar on the left of each pair is the conversation layer and the one on the right is the profile layer. The two bars stand on populations with different denominators. Rather than comparing heights, read only which way each pair tilts.
4.1What makes it look like a doubling
This is the point where a sentence about health information doubling from 21.31% to 41.11% wants to be written. The two values stand on different populations. A profile entry comes into being only when the memory feature pulls a fact out of a conversation. Conversations like that account for 7,051 of the 179,057, which is 3.94%. A sample assembled from conversations that contained facts will show a higher rate of health talk, and that is a selection result rather than an increase. Setting two layers side by side requires putting the denominator inside the sentence, and that is the first lesson of this table.
Something to add while denominators are in view. Recompute the risk-band rates of 14.03%, 3.27% and 3.62% over 179,057 and they become 14.17%, 3.31% and 3.65%. The denominator implied by the paper's values is about 1% larger than the total conversation count. Pre-filter figures may have been mixed in, but the paper offers no account, and the arithmetic above is our own. We use the values as the paper wrote them.
4.2One pipeline, two layers, labels that disagree
A longer look at the table turns up a place where the two layers do not add up. Profile entries carrying health talk number 2,899, while entries with a risk score of 2 or more number only 1,291. So 55.5% of the entries labelled as health talk, 1,608 of them, were simultaneously labelled no-risk. On the conversation layer this does not happen: there the 37,828 health conversations and the 37,828 conversations scoring 2 or more are exactly the same number. The same pipeline and the same model appear to have applied the labelling rule differently across the two layers. The paper does not explain the mismatch, and we are the ones pointing it out.
4.3The grade distribution's blind spot
Does that make the paper's claim wrong? No. What the paper calls risk was never a rate to begin with. Section 5.2.2 says so directly, and this passage is the centre of the paper.
Two points carry the argument. One is density. Once scattered cues gather on a single line, the number of people that line could describe, the anonymity set, shrinks. The risk-grade distribution cannot see that gathering. A grade looks at one line on its own and never counts lines piling up under one person. The other is permanence. Conversation feels to the user like something that flows past and stays where it happened, while a profile entry remains and rides along into later conversations. The scoring rubric is built that way: the prompt instructs the model to treat each detail such as occupation, age, location and language as an anchor that narrows the anonymity set.
Here the sister paper from section 2 walks back in. Whether the density really does single people out is the one thing this paper could not measure, and it has already been measured over the same 1,057 people. After keeping only conversations in which no message stated who the user was, a general-purpose model recovered age, gender and country at weighted F1 of 0.84, 0.90 and 0.88. So the precise statement becomes this. Evidence that profile storage raises the risk grade is not in this paper. Evidence that identity attributes can be recovered from the style and topics left in conversation exists over the same people. Grades being equal does not make the two datasets the same dataset.
A profile line has no slot for when
The paper's appendix carries six actual profile entries. All six begin with "User" and all six are present-tense state descriptions. The symptom-category example looks like this: this user has a persistent dry cough, runs a low-grade fever in the evenings and gets a little short of breath on ordinary walks. The stigma-category example records that the user is on outpatient medication and keeps track of the opening hours of a nearby clinic. The sociodemographic example has the user at seventy-four, minding grandchildren all day, living in a house with steep stairs.
The status of these sentences comes first. They are not what real users said. The authors synthetically substituted the details with Llama before release, and it is a six-row example table. It shows the shape of the schema rather than any measured distribution. And that substitution left a trace. The prompt instructed the model to swap place names "within the same regional context," and the examples that came out mention the Mayo Clinic and the Cleveland Clinic, a US insurer, rural Canada and Bogotá. That is outside the four countries under study. The sensitive details were erased and the context moved to other countries. This observation too is absent from the paper. It is a small case of what else comes along when a large language model is used to make a derivative.
Back to the schema itself, and something is missing from all six lines. When this was observed, how long it is taken to hold, what it was written from, and who wanted this line. The authors' sentence quoted in section 4 names that loss precisely: conversational nuance, temporal qualifiers and user intent are stripped away. Stripped of those, the line leaves behind a state. A cough asked about once last night becomes a symptom being suffered.
5.1What it means to say contextual integrity breaks
The name the paper gives this loss is a collapse of contextual integrity. Information staying within the norms of the setting where it was first shared is what contextual integrity means. When someone asks a chatbot about a symptom alone at night, the norm they assume is transience: this conversation stays in this window and ends there. The moment the content is stored as a state with the temporal qualifier gone, that assumption breaks. At that point the conversation stops being a one-off enquiry and becomes something closer to a clinical record.
5.2The vendor's remedy is an overwrite
The product knows that entries go stale. The June announcement explains its fix with a travel itinerary. Over time memory updates itself, so an entry saying a trip to Singapore is planned for July gets rewritten after the trip into an entry saying the user travelled to Singapore in July 2026. The direction matters: the remedy rewrites the entry rather than inscribing a time on it. No slot appears for when this was observed or how long it holds, and the previous state is quietly overwritten instead. Freshness rises and the path back does not survive. On the axis of auditability it moves toward disappearance. We read the direction that way; the paper does not.
Deletion is much the same. The help page says that removing something the chatbot might know entirely requires deleting "every source where it appears," and lists past conversations, archived conversations, files, memory summaries and connected apps together. It also says that turning memory off and back on can produce new entries from old conversations still in the history. Another sentence says logs of deleted saved entries may be retained for up to 30 days for safety and debugging. Deletion does not follow the derivative.
How profile entries are injected into later replies, whether they are used for training, and whether deletion propagates to derivatives are things this paper did not measure. What the paper measured stops at storage. The sentences above are descriptions in product documentation, not measurements. The piece on deleting agent memory, published here two months ago, found the owner of the line missing; the gap in the profile line is its time. Knowing the owner without the time leaves no way to set a retention period, and knowing the time without the owner leaves no way to receive a deletion request. Both are needed separately on the same entry.
The way to run this count again is gone
This audit was possible for a simple reason. Users could export their own data, the exported file contained the stored entries one by one, and so they could be counted from outside. The paper's limitations section records that this condition has gone. In the authors' account, in September 2026 the company moved the old per-entry memory feature off the default and introduced a "Memory Summary," and individual entries stopped being visible on the default screen. The option to revert to the old setting still exists, and that is what this paper looked at, but the text says that "to the best of our knowledge, export files no longer contain memory entries." The hedge is the authors' own and is carried over intact.
The product documentation points the same way. The help page says a memory summary "does not include everything ChatGPT remembers based on your conversations," and directs anyone who wants to know what it remembers to ask inside a conversation. Asked whether something missing from the summary is therefore absent from memory, it answers "not necessarily," adding that some details may not appear in the summary if they are judged less relevant or unsuitable for display on that screen. The same document notes that a route back to the old saved-entry list remains in settings. So the accurate statement is not that checking has become impossible. It is that the route of downloading a file and counting from outside has closed. This differs from the problem of conversations having nowhere to go when a companion service shuts down, covered here in July. That one is about a service disappearing. In this case the service carries on being updated while the audit window closes.
6.1The same company writes differently for hospital customers
The help documentation has a separate section for customers in regulated industries. It says that in healthcare products and in enterprise products with a regulated workspace attached, the improved memory feature is off by default, and beneath that sits a callout box carrying this instruction: the feature is not covered by a Business Associate Agreement under US health privacy law, so do not enter protected health information. The paper measured the result of 1,057 consumer users in India, Nigeria, Brazil and Pakistan putting their own health talk into exactly that feature. The point to draw here is not legality but location. The warning is written in one place and the measurement was taken in another.
6.2The law that speaks most clearly sits outside this sample
Is there a provision anywhere requiring that attributes derived automatically from conversation be handled in a particular way? With that question we fetched and read the statutes and judgments of the four sample countries and of the European Union in the original. No law-firm commentary was used. The result was not what we expected.
| Jurisdiction | Share of sample | How derived or inferred health attributes are handled |
|---|---|---|
| India | 456 (43.1%) | Across the full text of the Digital Personal Data Protection Act 2023, the words sensitive, profiling, inferred and special category appear not once. The §2(t) definition has a single category only, so health talk sits in the same box as a delivery address |
| Nigeria | 206 (19.5%) | Section 65 of the Nigeria Data Protection Act 2023 lists health status as sensitive personal data, and section 30(1) bars processing sensitive personal data before setting out nine exceptions, the first of which is consent for the specific purpose. Sections 27, 34 and 37 name automated processing, including profiling, and attach notification, objection and human intervention as rights. The definitions section, however, has no entry for profiling |
| Brazil | 205 (19.4%) | Article 5(II) of the LGPD lists data referring to health as sensitive personal data, and Article 11(I) requires consent given "in a specific and highlighted manner, for specific purposes" |
| Pakistan | 190 (18.0%) | There is no comprehensive statute. A May 2023 draft posted by the responsible ministry defines profiling as automated processing that uses personal data to analyse or predict attributes such as health, and lists health data as sensitive personal data |
| European Union | 0 | Two Grand Chamber judgments already reach derived and inferred health attributes. The October 2024 judgment held that such information counts as data concerning health "even where it is only with a certain degree of probability, and not with absolute certainty" |
Consulted directly on 17 September 2026: the full text of the Indian statute in the Gazette of India, the scanned Nigerian gazette together with a section-by-section edition, the original of Brazil's Lei nº 13.709/2018, the May 2023 draft published by Pakistan's ministry of information technology, and the full judgments in C-184/20 (1 August 2022) and C-21/23 (4 October 2024) on EUR-Lex. The word counts for the Indian act are ours, taken over the full text. The Nigerian gazette was read by character recognition and checked character by character against the section-by-section edition. Enforcement decisions were again not found. A provision existing and something having been decided under that provision are different matters, and this table holds only the first.
With those two rows in, the centre of gravity of the table shifts. Only one jurisdiction, India, is missing the vocabulary altogether, and the Nigerian act has the words it needs. It names automated processing including profiling in three sections, requires that the existence of such processing be disclosed in advance, and attaches a right to human intervention where a decision rests on automated processing alone. The regulator's 2025 implementation directive attaches a procedure as well. Processing that amounts to evaluation or scoring, and processing that involves sensitive data, carry a mandatory impact assessment that must be filed with the commission before processing begins. Both conditions appear, on their face, to overlap with the pipeline this paper audited. Overlapping and breaching are different claims, and this article writes only the first.
Pakistan catches the eye for a different reason. Two places in this table put the act of inferring or predicting health attributes directly into a sentence, the EU judgment and the Pakistani document, and the second is not a law. The draft's definition reads that profiling is automated processing that uses personal data to analyse or predict attributes concerning the data subject, health among them. It holds the pipeline this paper measured without altering a word, and it has been sitting in a ministry's document folder as a draft since May 2023. In the list of current-session legislation the National Assembly publishes, we did not find this bill. Set the four countries down one line each and it comes out like this. India lacks the word. Nigeria uses the word without defining it. Brazil requires consent given in a highlighted manner. Pakistan has the sentence still in a draft.
The two judgments have to be read together. The 2022 judgment held that information capable of disclosing something indirectly falls within the special categories, but that holding concerns sexual orientation. Inside the same judgment is a passage distinguishing "reveal" from "concerning," and health takes the latter word. The 2024 judgment fills in the health side. It held that when a pharmacy sells non-prescription medicines online, the name, delivery address and product details the customer enters constitute data concerning health, and that this holds even where it is only probable rather than certain that the medicines are intended for that customer.
Lay the two over each other and this reading follows (and the reading is ours). Under EU law, a derived line such as "user experiences a persistent dry cough" would most likely be treated as health data within the meaning of Article 9. Probable without being certain is enough, according to the Grand Chamber. The question then stops being whether the law reaches and becomes which of the Article 9(2) exceptions the processing stands on, and nobody has asked that yet. Meanwhile the only jurisdiction in this table with a holding that an inferred attribute is itself health information is the European Union, and not one of the paper's 1,057 people lives there. The most precise measurement and the clearest norm are in different countries. We do not write that down as a finding of breach. The paper performed no regulatory analysis either, and mentions comparing the Indian and Brazilian regimes only as future work.
6.3The control the paper recommends already ships next door
One of the paper's four design recommendations is per-category retention control, on the grounds that today's options amount to switching everything off or extracting without limit. That control already ships, off by default, in a competing product. Note first that the table below carries the wording of public documents. The authors write plainly that they could not find the corresponding entries in the export files of two other services and so could not compare. We measured nothing either. We read the public documentation as it stood on 17 September 2026.
| Product documentation | Saves without being asked | Entries visible one by one | Sensitive categories switchable on their own |
|---|---|---|---|
| ChatGPT | Yes | Partly. The summary does not hold everything, and a route back to the old list exists | Not in the documentation. Off by default only in regulated workspaces |
| Claude | Yes | Yes. Settings list them all by topic, and entries can be edited or deleted individually | Yes. Sensitive topics such as health are not stored by default, and switching the setting off is stated to delete sensitive entries already stored |
| Gemini | Yes | Not in the documentation. It suggests asking inside a conversation to find out | Not in the documentation |
Taken from each vendor's official help documentation, consulted directly on 17 September 2026. These are descriptions in documentation, not audit results. For reference, the paper's body says other vendors' memory architectures rely mainly on context the user manages directly, but its support for that is a separate paper about developer-facing agents, and the current text says the products store on their own. Rather than the paper being wrong, the text may have changed since the February 2026 collection window.
This table is not a ranking. Nobody has measured which is safer. Read it for one thing instead. A product where that design recommendation is already implemented sits elsewhere, and the product where the audit actually happened is the one whose audit window has closed. The documentation says all three vendors save without being asked. The difference lies not in whether they save but in whether what was saved can be checked from outside.
Why Pebblous Cares
A large part of what Pebblous sells is making derivatives from originals. Attaching labels, producing summaries, pulling attributes out and putting them into a table. This paper measured exactly that process. Out of an original called conversation, a line beginning with "User" was derived.
7.1The schema of a derivative makes claims the original never made
Conversation has tense. A sentence saying a cough started yesterday has a starting point, an open end, and a degree of confidence riding along with it. The derived line becomes: this user has a persistent dry cough. An observation gets promoted to an attribute, and this is not confined to medical conversation. An inspection note on one machine, the driving characteristics of one driver, a disposition tag on one customer. The derived fields handled every day in quality diagnostics have the same shape.
The observation in section 5 settles on top of that. When a derived field went stale, the remedy the product offered was rewriting the line. Managing by update raises freshness, and the previous state disappears with the path back. Data quality and auditability part company here. The two usually travel together, and a single update policy can send them in opposite directions.
7.2When the risk indicator holds still and the quality degrades
The most instructive thing in this paper is the table in section 4. The high-risk share was 3.62% on conversations and 3.63% on profile entries, practically the same, and by the score-2-or-higher measure the profile layer was actually lower. The shares taken by the symptom category and the stigmatised-condition category also shrank on the profile side. On the risk score alone, nothing happened. Yet in the same data the temporal qualifiers had gone, and cues that had been scattered were now gathered on one line.
Measure an original and a derivative with the same ruler and the grades come out equal. The structure changed, not the grade. How many cues one line holds, how long that line stays, whether that line can be traced back to the original. And section 4 already showed what those gathered cues make possible. The Nordic health data metadata audit, covered here earlier this month, measured how fully existing fields had been filled in. This case is one where the schema has no such field to begin with. The two problems call for different diagnostics. A completion rate can be counted, and what is absent from the design cannot be counted at all.
Third, when two layers are compared across different denominators, the comparison itself manufactures a claim. Health talk looks as though it doubled from 21% to 41%, but the second figure was measured over conversations that produced a profile entry, and those are 4% of all conversations. Drop that line from a quality report on derived data and a selection effect reads as a finding.
7.3Three things to check in practice now
Put the counts so far in front of a derived table and three places to check appear. One comes from the schema in section 5, one from the measures in section 4, and one from the audit route in section 6. All three follow directly from what this paper actually handled, and the third is also the reason the paper's method no longer runs today.
- Put four mandatory fields on every derived field. Observation time, validity horizon, evidence (which original it came from) and requesting party. That is a step ahead of the do-not-store flag the paper proposes. Without those four, no slot comes into being for attaching a retention period or receiving a correction request. And decide in the same place whether an update that overwrites keeps the previous value.
- Do not measure a derivative's risk with the original's ruler alone. Identical grade distributions still mean a different dataset if the number of cues per entry and the time the entry persists have changed. When two layers are compared, put the denominator inside the sentence.
- Write into the product requirements whether derived entries are included in export and deletion. Before it is a convenience feature, it is a question of verifiability. The reason this paper's method no longer runs is precisely that those entries dropped out of the export. Check as well whether deletion propagates to every source.
No prescription to stop deriving follows from any of this. Pulling context out of conversation and carrying it forward is a feature users want, and the paper's own design recommendations stop at visibility, per-category control and local processing. What remains is not a prohibition. It is fields, denominators and an audit route.
7.4Where Pebblous stands
This article leaves behind the shape of the fields missing from a derived-data specification. A data card today mostly carries the provenance and scale of the original, plus the derivation method. When the derived field was observed, how long it is taken to hold, and at whose request it was made are mostly not written down. And no amount of careful measurement of the risk-grade distribution will surface that omission.
The omission is not there because the work is hard. All this paper did was count 7,051 stored lines and look at the sentence in front of each. The hard part is not the measuring but the deciding to measure, and this paper demonstrates on its own case that the window closes once the product updates. For anyone whose work is designing the fields of derived data and diagnosing quality through those fields, and Pebblous stands there, that is the place these counts point to.
Korea has a case of similar shape. In last year's Iruda judgment, the court found that information allowing a specific individual to be inferred remained, even though the developer argued it had de-identified the data. The shape matches the tension in section 2. The mechanism differs, though. That case concerned raw conversations used to train a model, and this paper audited derived text entries stored by a product. It should be read as an analogy only.
Sections 1 through 6 carry what the researchers measured and what we verified ourselves in primary documents. The collision between de-identification and the high-risk measure, the observation that synthetic substitution relocated the regions, the reading that the audit window has closed, the implications of the uncited sister paper, the geography of the norm, and this section 7 are things the paper did not do. Please read them apart. And this article is not a finding of illegality against any product or any user. It is a measurement report on the derived data that consumer chatbot memory features create. The conversations here are records that participants donated with consent, and every vendor document and legal provision cited is the text as it stood on 17 September 2026. Text changes. Thank you for reading this far.
References
The numbers in this article come from three streams. Values from preprint 1 were carried over after checking the body, tables and appendices of the arXiv release against each other, and values that exist only inside a figure were not carried over. Product documentation, statutes and judgments were fetched and read in the original, and because their text changes, the access date sits alongside each. The Internet Archive had suspended service at the time of our lookup, so not a single snapshot of an earlier version could be kept.
The backbone of this report
- 1.S M Mehedi Zaman, Md Mozammel Hoque. "Vulnerabilities in Personalization: Assessing Health Privacy Risks in ChatGPT Logs and Memory." Preprint, arXiv:2609.14697v1, submitted 13 September 2026, cs.HC and cs.CY, CC BY 4.0. arXiv: 2609.14697 — the primary source for this article. The body, ten tables, appendices A through D and 41 bibliography entries were all checked. The arXiv Comments field is empty and some conference-template placeholders remain unsubstituted, so it was read as an unreviewed preprint.
- 2.S M Mehedi Zaman, Kiran Garimella. "Inferential Privacy Leakage in Anonymized Conversational AI Logs." Preprint, arXiv:2605.23820v1, submitted 22 May 2026, CC BY 4.0. arXiv: 2605.23820 — the sister paper from four months earlier, using the same donated corpus, the same ethics approval, the same 1,057 people and the same Llama-3.3-70B. The weighted F1 values of 0.84, 0.90 and 0.88 in sections 2 and 4 and the verbatim from its section 6.1 come from this sister paper, and the distinction between inference and memorisation in section 1 is its section 6.2. Paper 1 does not cite it (zero matches across the full body and bibliography). Its abstract and its section 6.1 describe the cohort filter differently, so the body carries only the conservative form.
Prior work and theory
- 3.Abhisek Dash et al. (2026). "The Algorithmic Self-Portrait: Deconstructing Memory in ChatGPT." Proceedings of the ACM Web Conference 2026 (WWW '26), 3471–3482. DOI 10.1145/3774904.3792671, arXiv:2602.01450 — the prior work referred to in section 3. Eighty users, 2,050 memory entries, 96% system-generated alone. It is also the original source of the documentation screenshots paper 1 reproduces in its appendix, taken in January 2026. It is a separate study with a sample more than thirteen times smaller, so the figures were not mixed.
- 4.Tawfiq Ammari, Michelle Chen, S. Zaman, Kiran Garimella (2025). arXiv:2505.24126 — used to establish the lineage of the corpus. The first author of paper 1 and a co-author of paper 2 both appear on it.
- 5.Chowdhury and Garimella. "How people use ChatGPT: conversation-level evidence from India, Nigeria, Brazil and Pakistan." — the work paper 1 gives as the source of its corpus. The bibliography entry has no venue and no URL, and the citation year is 2018, so it does not stand as given, and we could not find a published version either. Section 2.2 carries that state over as it is.
- 6.Helen Nissenbaum (2004). "Privacy as Contextual Integrity." — the source of the contextual integrity concept in section 5.1. The original was not consulted, so only the one-sentence definition is used and nothing is quoted verbatim.
Product documentation (all accessed 17 September 2026)
- 7.OpenAI. "Memory FAQ." help.openai.com/en/articles/8590148 — the verbatim in sections 3.1, 5.2 and 6 is taken from this page. The page's own stamp puts it at an early September 2026 revision. Direct requests were blocked, so the body was retrieved through a reader proxy. The February 2026 version could not be checked.
- 8.OpenAI. "Memory and new controls for ChatGPT." 13 February 2024. openai.com/index/memory-and-new-controls-for-chatgpt — the verbatim promise about sensitive categories in section 3.2 comes from here. The verb is "steer … away from" and "We're taking steps to" precedes it. The widely circulated "trained not to proactively remember" form is a paraphrase of that sentence and was not quoted.
- 9.OpenAI. "Dreaming: Better memory for a more helpful ChatGPT." 4 June 2026. openai.com/index/chatgpt-memory-dreaming — sections 3.3 and 5.2 quote this announcement. Counting sensitive, health, privacy, consent, opt out and export across the body returns zero for each (the three privacy matches are footer navigation), and that count is ours.
- 10.Anthropic. "Use Claude's chat search and memory to build on previous context." support.claude.com/en/articles/11817273 · Google. "Get personalization based on your past Gemini chats." support.google.com/gemini/answer/16598469 — the two rows in the section 6.3 table. Descriptions in documentation, not measurements.
Statutes, case law and statistics
- 11.India. The Digital Personal Data Protection Act, 2023 (Act No. 22 of 2023), full text in the Gazette of India — the source of the word counts and the §2(t) definition in section 6.2. The counting is ours.
- 12.Nigeria. Nigeria Data Protection Act, 2023, ss. 27(1)(g), 30(1), 34(1)(a)(vii)(viii), 37, 65 — the Nigeria row in section 6.2. The scanned gazette edition was read by character recognition and the quotations were then fixed by comparison against a section-by-section edition. The section 65 definitions list has no entry for profiling (zero in both editions).
- 13.Nigeria Data Protection Commission. "General Application and Implementation Directive (GAID) 2025," NDPC/NDP ACT-GAID/01/2025, Art. 28(3) and 28(9). ndpc.gov.ng — the list of circumstances triggering a mandatory impact assessment and the duty to file before processing begins come from here. It is an implementation directive issued under sections 1(a), 6(c), 61 and 62 of the Act.
- 14.Pakistan. Ministry of Information Technology and Telecommunication. "Draft of the Personal Data Protection Bill, 2023," §2(dd) and §2(kk). moitt.gov.pk/Legislations — the Pakistan row in section 6.2. It is a draft published by the ministry rather than a statute, and the most recent version in the ministry's document list is dated May 2023. In the list of current-session legislation published by the National Assembly, no enacted data protection act was found.
- 15.Brasil. Lei nº 13.709/2018 (LGPD), arts. 5º II, 11 I. planalto.gov.br — the Brazil row in section 6.2.
- 16.CJEU Grand Chamber, C-184/20 OT v Vyriausioji tarnybinės etikos komisija, 1 August 2022 · C-21/23 Lindenapotheke, 4 October 2024. Full judgments on EUR-Lex — the European Union row in section 6.2. The earlier judgment distinguishes "reveal" from "concerning" and its operative part is limited to sexual orientation, so the health side rests on the later judgment. Using both together is what prevents a misreading.
- 17.World Bank Open Data,
SH.MED.PHYS.ZSandSH.XPD.OOPC.CH.ZS(based on the WHO Global Health Expenditure Database) — the last two columns of the table in section 2.1. Queried directly through the API on 17 September 2026, with the latest non-missing year given per country. These values are not in the paper. We queried them ourselves. - 18.Seoul Eastern District Court, judgment of 12 June 2025, case 2021Gahap104007 — the analogy in section 7.4. That case concerned raw conversations used for model training, so the mechanism differs from the derived storage this paper audited.
Neighbouring pieces on the Pebblous blog
- 19.Memory Without a Name Tag Can't Be Erased — referenced in sections 1.3 and 5.2 when drawing the boundary. The boundary splits into two problems: recording the owner, and recording the time.
- 20.A Work Filter Cut Half the Conversations Out of AI Use Statistics — essential reading for section 2.2. This paper steps into the space that piece marked as out of frame.
- 21.Five Ways to Make a Chatbot Remember the Conversation and An AI trained on your old records misses your new illness more often — separated in section 1.3 as the implementation side and the model-weight side respectively.
- 22.US Teens Hid It From People and Confided Only in the AI, New York Times Seeks Sanctions Over OpenAI's Deleted Chatbot Logs, When an AI Companion Shuts Down, the Conversation Has Nowhere to Go, Not one Nordic health dataset meets the EU metadata profile — referenced one line each in sections 2.2, 3.2, 6 and 7.2.