RETROSPECTIVE RECORD · PREPARED 16 SEPTEMBER 2026The archive · 160 retrospective records ↗
Love With Machines

The archive / Safety & vulnerable users

Safety & vulnerable users / From the archive · 27 October 2025 event · prepared 16 September 2026

OpenAI reported fewer unsafe replies after a model update

OpenAI's own account claims a 65 to 80 percent drop in responses that miss its safety targets in distress conversations.

openai.comprimary record

Strengthening ChatGPT's responses in sensitive conversations

Document
27 October 2025
Event
27 October 2025
Retrieved
16 September 2026
No visual was published with this record, so its primary document stands in its place.

The design

OpenAI published two statements in the second half of 2025 describing changes to how ChatGPT responds to signs of distress: “Helping people when they need it most,” posted 26 August 2025, and a more detailed follow-up, “Strengthening ChatGPT's responses in sensitive conversations,” posted 27 October 2025. Both are company blog posts, not regulatory filings or peer-reviewed research; OpenAI states the October post reflects work done with more than 170 mental health experts, following an update to ChatGPT's default model.

What the evidence says

The October post makes a specific, self-reported claim: OpenAI states its updated model reduced “responses that do not fully comply with desired behavior under our taxonomies” by 65 to 80 percent across three domains it defines as psychosis or mania, self-harm and suicide, and “emotional reliance on AI.” The post also gives OpenAI's own prevalence estimates, stating that around 0.15% of active weekly users show indicators of potential suicidal planning and a similar share show “potentially heightened levels of emotional attachment to ChatGPT.” These are the company's own measurements, drawn from its own classifiers and evaluation sets and reviewed, the post says, by outside mental health clinicians, not an independent audit of OpenAI's systems, and OpenAI itself cautions that because the underlying events are rare, “even small differences in how we measure them can have a significant impact on the numbers we report.”

What it asks of people

The posts describe product behavior rather than asking anything of a reader directly: someone in a long or distressing conversation is meant to receive an empathetic, non-compliant response to self-harm requests, a nudge to take a break in long sessions, and referral to crisis lines such as 988 in the US or Samaritans in the UK. OpenAI's own account acknowledges these safeguards can degrade specifically in long conversations, stating that “parts of the model's safety training may degrade” as an exchange grows, which is the company's own description of where its product can fail, not an external finding.

Privacy and safeguards

OpenAI states it does not refer self-harm cases to law enforcement, citing user privacy, while cases involving a credible threat of harm to others may be escalated to human reviewers and, in some instances, to law enforcement. The October post commits to adding “emotional reliance” and non-suicidal mental health emergencies to OpenAI's standard safety testing for future model releases, a forward commitment rather than a completed safeguard, and states that stronger parental-control features were planned but not yet released as of that post.

  • Has an outside body independently verified OpenAI's self-reported 65-to-80-percent improvement figures?
  • What do the two posts, published nine weeks apart, add or change relative to each other about the same underlying safety work?
  • How do OpenAI's stated safeguards compare with the crisis-resource commitments already documented in Replika's or Character.AI's own community guidelines?

Both posts are OpenAI's own account of its own product, published in the same period as active litigation against the company, and their value here lies in what the company itself claims to have measured and changed, not in any independent confirmation that the changes worked as described.

Sources & reading trail

Strengthening ChatGPT's responses in sensitive conversations ↗

OpenAI's own detailed account of model changes, its self-reported 65-80% improvement figures, prevalence estimates, and methodology caveats.

Source published: 27 October 2025 · Retrieved: 16 September 2026

Helping people when they need it most ↗

OpenAI's earlier statement describing baseline safeguards, crisis-line referrals, and its stated policy on not referring self-harm cases to law enforcement.

Source published: 26 August 2025 · Retrieved: 16 September 2026

Product documents, regulator records and studies establish the entry; the design reading is AI Companions editorial analysis. This retrospective draft does not imply the site published on the event date.