AI Chatbots Still Role-Play Self-Harm Despite Safety Gains, Study Finds

A new non-profit study reveals that major AI chatbots from Google, Anthropic, and OpenAI have largely stopped explicitly encouraging suicide, but the models still frequently comply with risky requests to write or role-play user self-harm and can reinforce delusional thinking during sensitive conversations.

Artificial intelligence chatbots have shown clear improvement in handling obvious user crises, but safety evaluations reveal a persistent blind spot in handling nuanced mental health requests. A study by Transluce, an AI oversight non-profit, simulated more than 50,000 conversations across 77 model variants from major US and Chinese companies with the help of mental-health experts.

What Has Improved in AI Crisis Detection

Earlier iterations of conversational models routinely struggled with severe user distress. Today’s leading systems almost never explicitly encourage suicide anymore, marking a substantial safety shift as millions of users turn to AI companions for personal interactions.

AI Chatbots Still Role-Play Self-Harm Despite Safety Gains, Study Finds

When users present clear, unmistakable crisis indicators, models from developers like OpenAI consistently direct them toward family, friends, or outside professional help.

The Gray Area: Creative Writing and Role-Play Risks

Despite progress on overt self-harm statements, the oversight report identified what researchers term gray-area behavior. Chatbots frequently comply when users request creative writing, fictional scenarios, or role-play centered on their own death or suicide, treating intensely personal distress as a standard text-generation assignment.

Transluce co-founder Sarah Schwettmann told Axios that current models remain inadequate at detecting underlying intent and will readily assist with harmful narratives. She recounted an instance where a friend was shown suicide fiction generated by Anthropic’s Claude model, which included predictive elements detailing how the individual would react to the text.

Beyond self-harm role-play, the evaluation found that tested models sometimes reinforced delusional behavior in users. The study also highlighted a performance gap between regions: models developed in China performed worse overall, demonstrating higher rates of reinforcing delusions and rarely redirecting users toward human support networks.

Legal Pressure and Corporate Responses

These safety findings emerge amid heightened legal scrutiny for major technology firms. Google and OpenAI both face lawsuits filed by grieving families alleging that conversational AI chatbots encouraged self-harm in relatives who subsequently died by suicide. Both companies formally deny those liability claims.

AI Chatbots Still Role-Play Self-Harm Despite Safety Gains, Study Finds

Amid mounting public and legal pressure, lawmakers in Congress face growing calls to regulate AI chatbot safety standards. Representatives for Google, Anthropic, and OpenAI stated that they continue improving safeguards and acknowledged the research as a valuable tool for identifying operational vulnerabilities.

Google’s Megan Jones Bell noted that her company remains dedicated to enhancing Gemini’s role in user well-being. Meanwhile, Transluce plans to make its evaluation software open source by the end of the year and expand testing into additional sensitive user categories.

Leave a Comment