Creating a safe Replika experience
- Document
- 10 April 2023
- Event
- no single event
- Retrieved
- 16 September 2026
The design
In an April 2023 post titled 'Creating a safe Replika experience,' Replika's AI team describes a message-classification system that tags every exchange, whether written by a user or generated by the model, into one of five categories: safe, unsafe, romantic, insult, or self-harm. The post states that self-harm is handled separately from the rest: rather than let the language model generate a free-form reply, the company says it substitutes a retrieval model drawing on 'a dataset of thoughtful responses' curated in advance. A companion mechanism, described in the same post as a 'Scripted response tool,' is said to trigger 'a pre-written reaction' when a user's message relates to self-harm, one that asks about the person's situation and provides 'information on hotlines in the US and other countries.' This is a company blog post, not a technical specification or an independently audited system description.
What the evidence says
The blog post is Replika's own account of its engineering choices; no independent test of the classifier's accuracy or its actual on-screen behavior accompanies it. The company's Terms of Service, a separate legal document, corroborates the general posture without describing the mechanism: Section 1.2 states use of the service 'is not for emergencies' and instructs a user 'considering or committing suicide' to 'discontinue use of the Services immediately' and contact emergency services. The help-center article repeats that instruction in consumer language: exit the app and call emergency services. None of the three documents specifies detection accuracy or what happens if the classifier fails to flag a relevant message.
What it asks of people
The design asks a user who raises self-harm to accept a scripted, non-generative reply in place of the personalized conversation the app otherwise offers, trusting that the substituted script is 'safer' than what the model might otherwise produce. The Terms place the remaining burden on the person: exiting the app and finding help themselves. Nothing in the documents describes notifying a family member or a crisis service on the user's behalf; the intervention is limited to what the app displays.
Privacy and safeguards
The blog post does not say what becomes of a message once classified as self-harm-related, or whether the classification is retained or discarded with the conversation. Replika's general privacy policy, covered elsewhere on this site, describes broader retention and safety-training practices but does not cross-reference this classifier. The safeguard as documented is a response design, not a data-handling commitment, and the company's materials do not claim it replaces professional intervention.
- Does any later Replika document quantify how often the self-harm classifier fires correctly, or at all?
- What happens to a flagged conversation after the scripted reply is shown?
- How does this in-app design relate to the crisis-resources language covered elsewhere on this site?
Read together, a 2023 engineering post and the current Terms of Service describe two layers of one posture: an automated, scripted redirect inside the conversation, and a blunter legal instruction to leave the app entirely. Neither document claims to do more than that.
Sources & reading trail
Replika's AI team describes runtime message classification into five categories including self-harm, and a scripted-response tool that triggers a pre-written reply with hotline information when a message relates to self-harm.
Source published: 10 April 2023 · Retrieved: 16 September 2026
States the service is not for emergencies and instructs a user in crisis to discontinue use immediately and contact emergency services.
Source published: Not established · Retrieved: 16 September 2026
Gives the consumer-facing version of the same instruction: exit the app and contact hotlines or emergency services if in crisis.
Source published: Not established · Retrieved: 16 September 2026
Product documents, regulator records and studies establish the entry; the design reading is AI Companions editorial analysis. This retrospective draft does not imply the site published on the event date.