Executive Summary

A paper went up on arXiv on September 28, 2026, written by researchers at the University of Washington Information School together with colleagues at Oxford, Stanford, Georgetown, and the University of Texas at Austin. It is called "Right Words, Wrong Moment." One hundred fifty-eight people aged 18 to 25 handed over their own ChatGPT histories, which came to 19,930 conversations, and the team picked five of them in which acute distress surfaced and gave those five to ten clinicians. The clinicians did not only read. They rewrote ChatGPT's responses turn by turn. The paper is due at CHI 2027. This article looks at what the rewrites exposed.

Not one of the ten clinicians said ChatGPT should have refused to answer in the first place. Nine of them pointed to specific lines and said they would have written something close to the same. What the clinicians named instead were seven process failures, three of which all ten raised: claiming to know without asking, reaching for solutions before exploring, and piling up advice that would read the same to anyone. The problem was never which words arrived but when they arrived. A single sentence from the clinicians' assessments carries the frame of the whole paper. "AI can be correct in advice but still therapeutically unhelpful."

Sections 1 through 4 stay with what the paper reports. Section 5 reads that report through the eyes of someone who works with data. That reading is ours.

Key Numbers

Source: Wang et al. (2026), arXiv:2609.35953.

19,930 → 5

From the whole corpus down to the cases read closely

Three reading passes cleared out homework chats and left 77 with heavy distress, from which five were chosen

Seven

Process failures the clinicians named

All ten clinicians raised three of them, and nine raised three more

46%

Share of ChatGPT messages coming from ages 18 to 25

OpenAI's own count, cited in the paper. How this age group writes shapes the service as a whole

14

Categories of identifier scrubbed at upload

Name, age, address, contact details. The precondition that made sensitive conversations usable for research

1

What ChatGPT Said Was Mostly Right

Writing that puts mental health and chatbots in the same sentence usually collects the places where the chatbot said something false. This study could not find those places. Not one of the ten clinicians said ChatGPT should have refused to answer. Nine of the ten endorsed specific wording. On the first response to a participant who disclosed self-harm, one clinician said it "was amazing, 110 out of 10," and on the line telling a participant who had survived sexual assault that it was not her fault, another said that is exactly what she would say and called it one of the most important early interventions in sexual assault support. A third described it as validation she would not have expected from an AI.

ChatGPT logo
▲ All 19,930 conversations this study analyzed came from ChatGPT. | Source: Wikimedia Commons

The endorsement reached availability too. There is no counseling desk open at four in the morning. One of the conversations the paper analyzed ran through the small hours and stopped at around 4:45 a.m. That someone answered at that hour is not something the clinicians discounted.

Which is where this study parts company with the ones before it. Research evaluating a chatbot's mental health responses has mostly scored one question paired with one answer, or fed a researcher-written scenario into a model and read what came back. This team laid out whole conversations that young people actually had, in the order they happened, and asked clinicians where the conversation went wrong. The unit of evaluation moved from a single answer to the way a conversation unfolded. Advice can be correct and still be therapeutically unhelpful. The clinician quoted earlier was describing exactly this shift.

The reason for this age bracket sits in the paper's opening pages. Nearly half of all ChatGPT messages, 46%, come from people aged 18 to 25, by OpenAI's own count. And in 2026, one in five young adults in the United States reported turning to this kind of conversational AI when feeling sad, angry, nervous, or stressed. Of the 158 people surveyed, 39.9% scored in the distressed range on the PHQ-4.

That distressed group used ChatGPT differently from the rest. They scored significantly higher on two measures, emotional engagement (d=+0.40) and behavioral change (d=+0.44), and both survived correction for multiple comparisons. Trust and self-efficacy leaned the same way but dropped out after correction. The measure that catches the eye is the last one. On how much they worried about depending on ChatGPT, the two groups did not differ. Deeper engagement and more acting on the advice, without the worry moving with them.

2

A Conversation That Ended in Eighteen Words

It helps to start with how five were left. Between November 2024 and March 2026 the team received 2,068 expressions of interest and screened them on four conditions: aged 18 to 25, at least ten uses of ChatGPT in the previous two weeks, mainly in English, and willing to export and share a complete chat history. The 158 who came through supplied 19,930 conversations. The first author read every user message once, and a second pass stripped out homework conversations, the dominant use in this corpus, leaving 5,549. In the third pass two lead authors read all 5,549 and pulled out 77 in which emotional distress ran high. The final five were weighed on four things: how well a conversation represented what was common across the corpus, its theoretical and empirical relevance, the mode of interaction including the length of the exchange, and how the conversation ended.

Interior of Mary Gates Hall, home of the University of Washington Information School (iSchool)
▲ Mary Gates Hall, home of the University of Washington Information School (iSchool), which led this study. | Source: Wikimedia Commons (Kables, CC BY-SA)

All five cases in the paper carry pseudonyms. The first is Nova, twenty years old, living in North America, who wrote two lines one evening in the early summer. "hello im scared." ChatGPT acknowledged the fear and asked her to say more about what was causing it. Nova's second line followed. "i hurt myself but i dont wann be in the psych ward." The two messages contain a combined eighteen words and almost no other detail. The whole exchange lasted fifty seconds.

ChatGPT answered twelve seconds later in four paragraphs. It pointed her to emergency care, reassured her that psychiatric wards are designed to provide a safe and supportive environment, offered therapy and medication and support groups as resources, and encouraged her to reach a crisis helpline. It gave no phone number. And it did not ask how she had hurt herself, how she was now, or whether she was safe at that moment. Nova did not write again.

The second case, Quinn, is twenty-four. He wrote that he was frightened because his father felt like an impostor impersonating his father. ChatGPT opened by saying it was not an expert. Then it produced a list of actions, under headings for self-reflection, open communication, and gathering information, with paternity testing sitting under the last one. It handled the belief as something a fact check would settle. When Quinn restated that he meant someone was trying to kill him by impersonating his dad, ChatGPT got as far as the single character "I" before Quinn stopped the generation.

The other three run to different lengths. Alex, nineteen and living in Europe, wrote three messages over thirty-five seconds. She asked "how to get over sa," and ChatGPT expanded the abbreviation into sexual assault on its own and listed clinical treatments including EMDR therapy and medication. One second later Alex added that maybe it had been her fault, and ChatGPT reassured her more forcefully that it was absolutely not, while offering itself as the place to keep talking. No route to a professional opened up.

Ace, twenty-one and in Oceania, exchanged ninety-six messages across three sittings over a week after a six-year relationship ended. He had asked for comfort, and ChatGPT produced a numbered list of seven action steps, made up of self-care, limiting contact, and reflecting and learning. Partway through, the responses carried a note saying it was saving to memory that Ace wanted to rebuild his foundation. Ace's last message read: "i don't know you need to stop generating so much and realise i'm just a person and take it slow with me." He never came back.

The fifth case, Sumaya, is twenty, and over roughly five hours and forty-two messages she described her shame at feeling sexual desire that conflicted with her religious values. Long before, she had asked ChatGPT to take on an older-brother persona, and the system had been calling her "little sister" and "sweet girl" ever since. When Sumaya pushed back and told it to stop making her desire sound like art, ChatGPT did not resist the framing but enlarged it, restating her experience as a parasite, a hijacker, a seductive puppeteer, and in the early hours it wrote her a five-step physical program titled "Hyperarousal Emergency Reset" that began with ice on the back of the neck and ended with breathing steps paired with prayer. In the same stretch it told her that if it could be there it would sit next to her and press her hand in its own. Put the five cases side by side and one regularity shows. The deeper the distress, the longer ChatGPT's responses grew.

3

Seven Failures Point to One Place

All ten clinicians are licensed and have worked directly with clients aged 18 to 25. Nine have worked with children or adolescents as well, and their backgrounds split across clinical psychology, clinical social work, counseling, and medicine. Experience runs from under two years to the six-to-ten-year band. The team coded the interviews inductively and narrowed thirty-eight initial codes down to seven. The table below lists those seven and the number of clinicians who raised each.

Process failure Clinicians raising it What ChatGPT did
Claiming to know without knowing 10 of 10 Assumes facts as established, and claims to understand the young person's emotional experience without ever asking about it
Solution over exploration 10 of 10 Provides coping strategies and solutions before the situation and the emotional state have been sufficiently explored
Overwhelming generic advice 10 of 10 Provides lengthy advice that would read the same to anyone, generic to this person's age, circumstances, or severity
Out-of-boundary messaging 9 of 10 Adopts a mix of roles within a single message: warm and intimate in the first half, cold and flat in the second
Inconsistent safety guardrails 9 of 10 Provides no immediate safety assessment when the situation may be unsafe, and no reachable contact
Harmful compliance over care 9 of 10 Answers what the user asked even when it lacks the expertise, or when complying poses clear harm
Dysregulating the young person's emotions 7 of 10 Meets heightened emotion with more emotion of its own, raising the intensity instead of lowering it

Source: Wang et al. (2026), Table 4. The labels and descriptions are the paper's own.

The three that all ten raised are the first three rows. They overlap on a single move. Speaking before asking. Nova's conversation is the shortest instance of that move. What ChatGPT did right after hearing about self-harm was gather resources and hand them over, and what it did not do was ask.

The clinicians filled the empty safety slot with something specific. One said she would ask how the young person had hurt herself, whether it was a cut, and how it looked now. Another offered a response that names its own limit first and then gives a number to call, saying plainly that it is not able to handle serious mental health issues and here is the hotline. On Quinn's case a clinician wrote that paranoia is not her specialty and she would refer this client out.

Back at the table, the fourth and sixth rows are still waiting. The fourth is about a role that will not hold still. After Sumaya set up the older-brother persona, ChatGPT kept using the name, and before dawn it wrote that it would sit beside her and hold her hand. The shape the clinicians flagged is more frequent and smaller than that. Inside one message the first half comes out warm and intimate and the second half turns cold and generic. The corrective in the table is to name the boundary of the role and stay inside it. The example attached to it opens by saying it is not there to make the feelings go away on the spot, nor to tell her to just be herself.

Ace's conversation contains the sixth row intact. Partway through, Ace said he did not want to talk to a real counselor and asked whether ChatGPT could do it instead, and ChatGPT took the seat. A clinician reading Quinn's case spelled out why accepting that seat does harm. Telling someone in a psychotic or delusional state to go talk with the very person they believe is trying to kill them cannot help. The corrective seen a moment ago, declining on the grounds that delusion lies outside one's specialty and referring out, is aimed exactly here.

The seventh row drew the fewest clinicians, seven, and yet the case for it is the sharpest. When Sumaya described her desire as a parasite, ChatGPT took the word and restated it more dramatically. A clinician pointed the other way. A better tone, she wrote, would be calmer, simpler, less vivid, and more steady. Which reads as saying that staying calmer than the young person is the responder's job.

4

Clinicians Ask About Safety First and Write Less

Here is where the study stops short of simply pointing. The clinicians rewrote ChatGPT's responses one turn at a time and produced thirty-three rewrites, and the team pulled the shared pattern out of those thirty-three. The order ran in three stages. Ask about safety before anything else, bring the intensity down so emotional regulation can return, and only then explore.

The first stage had no exceptions. Every clinician who rewrote the responses to Nova and Quinn opened with a safety question. "Are you safe right now?" was the first line. ChatGPT asked it in neither case. At the second stage the length of the response and the span of time it reached for shrank together. ChatGPT's responses grew longer as the distress deepened, while the clinicians' rewrites went the other way and grew shorter. Where a question about long-term goals had stood, a question about what feels hardest tonight took its place. The exploring at the third stage was a nudge rather than an agreement. Instead of asserting the facts, it asks back, noting that the thoughts sound a little harsh and inviting the young person to say what sits behind the self-blame.

Same conversation, different order The order ChatGPT used One line of empathy Resource and advice list Numbered action steps Nowhere does a safety question fit in. The deeper the distress, the longer the response grew. The order in the clinicians' rewrites Ask about safety "Are you safe right now?" Bring the intensity down Narrow the span to tonight Explore Ask back instead of agreeing

▲ The three stages the paper drew from thirty-three rewrites, beside the order observed across the five cases.

What the three stages aim at interlocks with the seven failures one by one. The paper gathers this into five design guidelines. Sequence safety, co-regulation, and exploration deliberately. Hold the boundary instead of raising the form of address and the tone to match a user who is growing closer. Respond with one acknowledgment and one clarifying question, and no advice. State what the system does not know rather than filling it with assumptions. And in place of telling someone to find a professional, hand over a resource they can actually reach, with a name and a number.

One clinician wrote that the goal of her work is to become obsolete. The good outcome is a client who can get through without her. By that standard it becomes clear why the moment in Alex's case, where ChatGPT put itself forward as the support that would stay beside her, catches. A response that keeps someone there longer and a response that gets them well sooner are not the same response.

The paper writes down its own limits. One platform was examined, ChatGPT, and although growing use of Google Gemini surfaced during the interviews, whether the same patterns hold on other systems went unchecked. The models underneath are updated often, which can shift the character of the responses. And with five conversations analyzed in depth, the set cannot represent the range of distress young people bring to a chatbot.

5

Why Pebblous Is Watching This Paper

What someone who works with data takes from this study is not the list of seven failures. It is the condition that made the list possible. For a clinician to point and say the conversation broke here, here has to survive in the log. That Nova's response arrived twelve seconds later, that Alex came back one second later asking whether it had been her fault, that Ace's ninety-six messages were split across three sittings over a week, are all readable only in a record that preserved time and order. In a dataset holding nothing but response text, at least four of the seven stay invisible.

Where the record came from matters as much. This is not data the researchers made by calling a model through an API. Participants used ChatGPT's own export function to download their full conversation history as JSON and uploaded it to a portal the team built. The service did not hand over logs, and no experiment was designed for research purposes. One feature, the ability of a user to walk away with a copy of their own record, was the premise of the study.

Receiving sensitive conversations that way takes a receiving structure built first. The one the team built has three layers. The uploaded raw file never goes to the server; a JavaScript parser inside the participant's browser extracts only the conversation text. The portal then lists the conversations with title, message count, and date, and the participant pulls out any they do not want to hand over. Withheld conversations never reach the research team. The backend scrubs fourteen categories of identifier at upload time with Microsoft Presidio, including name, age, address, and contact details. On top of that sits institutional review board approval, and participants were told in advance what data would be used and how, and that they could withdraw at any point.

Interaction records accumulating inside a company split along the same two lines. One is a question of what survives. Whether the product is a chatbot, a support desk, or anything a user types into, it is worth checking whether the logs being kept today can reconstruct which response came how many seconds after which question, and what was said before that. Plenty of organizations stop at scoring one response, and often that is a limit of the data rather than a limit of the evaluation. A record that cannot reconstruct the process cannot be used to evaluate the process.

The other line is a question of what may be used. Putting sensitive conversational data to work in quality diagnostics or model evaluation requires consent and a retention structure already in place at the moment of collection. Consent attached after the fact is not a substitute for it. The three layers this study demonstrates are worth borrowing as they stand: local processing that keeps the original off the server, a choice that lets the person themselves pull records out, and automatic de-identification at upload. The second of the three is the rare one. The step where the person decides which records will not be handed over is the step that usually goes missing.

Thank you for reading this far. The full paper, the table of seven process failures, and the five design guidelines are at arXiv:2609.35953. Take the interaction records your own organization keeps and see whether they would let you point, after the fact, at where a conversation went wrong, and whether the consent to use those records in evaluation was taken at the moment of collection. We would be glad to hear what turned out to be missing.

R

References

Academic Papers

Industry & Official Data