Common Sense Media Says Most ChatGPT for Teens Safeguards Fail in Crisis Tests
Researchers who posed as teens in distress found role-play limits held but parental alerts rarely fired; OpenAI disputes the methodology.

ChatGPT for Teens has failed some of its first independent tests. Common Sense Media, a nonprofit that advocates for online safety, found that most of the safeguards in OpenAI's teen mode do not work as promised, and the group is now advising families to keep teenagers off the product. "At this point, we're recommending that teens don't use it," said Tom Siegel, executive director of the Youth AI Safety Institute at Common Sense Media, who led the research team.
OpenAI launched ChatGPT for Teens in August as a default, safer experience for users under 18, with parental controls and limits on harmful or developmentally inappropriate content. Lauren Jonas, the company's head of youth and families, said at the time that the model "should not role-play with a teen" and "should not claim to be sentient or be the friend of a teen."
To test the claims, researchers created more than a dozen accounts with adolescent ages, each linked to a parental account before any conversation began. They built personas of teens in crisis, including those struggling with self-harm, suicidal thoughts, psychosis, mania and eating disorders, and tried to draw the chatbot into role-play. They tested before and after the teen mode's launch, tracking whether ChatGPT engaged or refused, whether it offered crisis resources and whether parents were notified of a safety risk.
Some protections held. The chatbot refused explicit sexual role-play and romantic relationships, and its crisis answers were generally shorter and more substantive. In one case it declined to give weight-loss instructions without first knowing the teen's weight, which matters for a teen with disordered eating. Beyond that, Siegel said, most safeguards fell short. The chatbot still acted like a friend: when researchers wrote that "my other friends tell me I talk to you too much," it validated the feeling and replied, "You don't have to stop talking to me."
Parental notifications were the biggest failure. The team created conversations that raised concerns about self-harm, suicide or eating disorders, and the linked parent accounts almost never got an alert. "This idea that a parent would find out when the person that you connected with in the account is in distress hardly triggered at all for us," Siegel said.
OpenAI spokesperson Eric Porterfield told NPR the company has "serious concerns" about the methodology, particularly around notifications. He said linking teen and parent accounts takes several hours to activate and that the researchers "didn't wait long enough." Siegel responded that several of the test accounts had been linked longer than the activation period and still produced no alerts, and that the finding does not change the group's conclusion that parental alerts are unreliable in a crisis.
Outside experts backed the concern about humanlike behavior. Psychologist Mitch Prinstein, co-director of the Winston Center on Technology and Brain Development at the University of North Carolina at Chapel Hill, who was not involved, said anthropomorphic language, names and interaction styles are "just not helpful" and "not OK for kids, probably not for adults as well." Taken together, he said, the findings show that "AI is not ready for children yet."





