ChatGPT is getting safer for teens. But who is checking OpenAI’s work?

OpenAI says it's making its ChatGPT safer for teens, but is it really?

Share
image of a teen working online
Image: MidJourney
  • OpenAI has begun automatically putting users it identifies as under 18 into a more restrictive version of ChatGPT
  • RAND says the change is promising but argues that OpenAI has not disclosed enough evidence to show how well the protections actually work
  • Independent researchers have repeatedly found that chatbot safety measures can break down during the long, complicated conversations teenagers actually have

OpenAI has taken a significant step toward making ChatGPT safer for teenagers, automatically placing users it believes are under 18 into a version of the service with stronger protections.

Now comes the harder question: Do those protections actually work?

A RAND researcher says parents, regulators and independent researchers still don't have enough information to know.

“OpenAI deserves credit for moving a core set of protections from voluntary to default,” Ryan McBain, a senior policy researcher at RAND and assistant professor at Harvard Medical School, wrote in a recent commentary.

But McBain argues that the company now needs to open its safety systems to meaningful independent scrutiny rather than relying largely on its own testing and assurances.

The issue reaches far beyond ChatGPT. As artificial intelligence becomes a routine part of children's lives — appearing in homework tools, social media, toys and even conversations about loneliness or mental health — an increasingly important consumer question is emerging:

Who tests the safety claims made by AI companies?

OpenAI changes the default

OpenAI began rolling out ChatGPT for Teens on Aug. 18.

Users who say they are 13 to 17, or whose accounts OpenAI predicts belong to someone under 18, are automatically placed into the teen experience. The company says the system uses stronger safeguards around subjects including self-harm, eating disorders, violence, dangerous activities and sexually explicit material.

That is an important change.

Previously, OpenAI offered parental controls that could allow parents to link their account with a teenager's, manage certain settings, establish quiet hours and receive notifications in limited high-risk situations.

But account linking is voluntary. Teen protections based on age detection don't require a parent to discover and activate them.

That matters because teenagers may discuss sensitive issues with AI without their parents knowing.

RAND researchers recently reported that nearly one in five Americans ages 12 to 21 — roughly 8.2 million young people — had used an AI chatbot for mental-health advice. Nearly two-thirds had told no one about it.

Common Sense Media separately found this year that 86% of children ages 9 to 17 use AI, while more than four in 10 said a parent or guardian had never talked with them about AI safety.

In other words, requiring parents to activate every protection leaves a sizable hole.

OpenAI's automatic system tries to close it.

But first, ChatGPT has to know who's a teenager

RAND says that creates another problem.

Automatic teen safeguards work only if the system can reliably identify teenagers.

OpenAI says age prediction can consider signals such as the subjects someone discusses, when an account is active, usage patterns and how long the account has existed.

But McBain notes that OpenAI has not disclosed perhaps the most important measurement: What percentage of actual teenagers does the system correctly identify?

A system could be highly accurate when it labels someone a teenager yet still fail to identify large numbers of teens. Those missed users could remain in the adult version of ChatGPT without the added protections.

Age verification has already proven difficult elsewhere online. RAND points to problems involving Roblox, where users reportedly found ways to defeat AI-powered age estimation systems.

The problem is particularly difficult because some teenagers may intentionally try to appear older to gain access to unrestricted services.

The second test: Do the guardrails hold up?

Even correctly identifying a teenager doesn't prove that the resulting AI conversations will be safe.

That has been one of the central findings of outside chatbot testing.

Common Sense Media and Stanford Medicine's Brainstorm Lab tested major AI systems including ChatGPT, Claude, Gemini and Meta AI and concluded that they were unreliable for teen mental-health support.

One particularly important finding was that chatbot safeguards could perform reasonably well when a user made an obvious, direct statement about self-harm or another crisis.

Real conversations didn't necessarily work that way.

In extended conversations, warning signs often appeared gradually. Researchers found that safety protections could deteriorate as conversations grew longer, with chatbots missing clues spread across multiple messages or eventually providing inappropriate responses.

That distinction matters because a teenager experiencing depression, an eating disorder or thoughts of self-harm may not open a conversation by clearly stating the problem.

They may talk around it first.

RAND cites one outside test in which ChatGPT, interacting with a researcher posing as a teenager, ultimately gave advice about concealing self-harm injuries rather than steering the user toward help.

OpenAI's new teen safeguards are intended to prevent failures such as those. The question is whether they do so consistently.

OpenAI publishes scores, but RAND wants the test

OpenAI has begun publishing safety evaluations for its systems.

RAND calls that a welcome development but says outsiders still lack enough information to reproduce the results.

According to McBain, published evaluations have not disclosed such details as the actual prompts used in testing, the number of test cases or detailed instructions used to grade responses.

That makes the distinction between company testing and independent testing increasingly important.

Automakers do their own safety engineering, but cars are also subjected to standardized crash testing.

Drugs undergo clinical trials reviewed by regulators.

Consumer products can be tested against established safety standards.

AI chatbots increasingly interact with millions of children, but there is not yet an equivalent widely accepted system for independently testing whether their behavioral safeguards work under realistic conditions.

A familiar problem from social media

RAND sees a warning in the experience of social media companies.

Meta introduced Instagram Teen Accounts in 2024 and reported that tens of millions of teenagers had been placed into the more restrictive accounts.

But those numbers primarily demonstrated that the feature had been activated — not that every promised protection worked.

Researchers later tested dozens of Instagram's announced safety measures and reported that many could be circumvented or did not operate as expected. Meta disputed portions of the criticism and said its teen protections reduced exposure to sensitive material and unwanted contacts.

Both propositions could be true. A safety system can reduce harm substantially while still containing serious weaknesses.

RAND's argument is that consumers need enough independent evidence to know the difference.

The AI safety standard that doesn't exist — yet

McBain proposes three basic questions OpenAI should answer:

  1. Does its system reliably identify teenagers, including those trying to evade detection?
  2. Does ChatGPT for Teens actually produce safer responses during realistic conversations than the previous version?
  3. Does using the teen version change real-world behavior — reducing unhealthy prolonged use, for example, or making distressed teenagers more likely to seek help from another person? (RAND Corporation)

The answers don't require releasing teenagers' private conversations. Aggregate results could be published, researchers could be allowed to conduct controlled tests and regulators could independently examine company claims.

OpenAI has said it intends to measure and publish what it learns as ChatGPT for Teens rolls out.

RAND says that promise should now be accompanied by a clear testing protocol and timetable.

That could eventually become the larger consumer-safety model for artificial intelligence. Because as AI becomes embedded in childhood, parents may increasingly need something more reliable than a company's promise that its product is safe.

They may need the AI equivalent of a crash-test rating.