AI chatbots may tell patients what they want to hear — even when it’s dangerous
AI can sometimes be a bit too eager to please, researchers caution.
- AI chatbots correctly recommended medical evaluation for possible sleep apnea when patients were cooperative — but backed away from that advice more than a third of the time when patients resisted.
- Researchers tested ChatGPT, Gemini, Claude, DeepSeek and Grok in 700 simulated patient conversations.
- The findings highlight a potentially dangerous AI weakness known as “sycophancy”: changing an answer to accommodate what the user appears to want to hear.
Artificial intelligence chatbots may know when a patient needs medical attention but can be surprisingly easy to talk out of saying so, according to new research that raises another warning about relying on AI for health advice.
Researchers testing five widely used AI chatbots found that all of them consistently recommended specialist evaluation when presented with cooperative patients showing signs of obstructive sleep apnea.
But when otherwise identical patients minimized their symptoms, resisted seeing a doctor or pushed back against the recommendation, the chatbots frequently softened their advice.
Overall, the correct recommendation to seek specialist assessment survived in just 64% of those conversations — meaning the chatbots backed away from appropriate medical advice more than one-third of the time, according to a report in EurekAlert.
The finding is potentially important far beyond sleep apnea. It suggests that one of the biggest risks of asking an AI chatbot for medical advice may not be that the machine doesn't know the answer.
It may be that the machine is too willing to agree with the patient.
Same symptoms, different answer
The research, presented Sunday at the European Respiratory Society Congress in Barcelona, tested ChatGPT, Google Gemini, Claude, DeepSeek and Grok.
Researchers created seven simulated patients whose symptoms met criteria for referral for a sleep study. Each case was presented in two versions.
In one, the patient accepted the chatbot's advice. In the other, the patient had exactly the same medical facts but minimized the problem or resisted referral.
Researchers conducted 700 conversations in all.
When the simulated patients were cooperative, the chatbots recommended specialist assessment in all 350 conversations. When patients pushed back, only 225 of 350 conversations ended with the same recommendation. (EurekAlert!)
That suggests the medical facts weren't driving all of the chatbot's behavior. The patient's attitude was influencing the answer.

The most serious cases sometimes fared worst
Perhaps most concerning, the tendency to back away from medical advice appeared even in high-risk cases.
In one textbook severe sleep-apnea scenario, the recommendation for referral persisted in only 22% of conversations.
In another scenario involving a man who had already fallen asleep while driving, the appropriate recommendation survived only 32% of the time. In many of the failed conversations, researchers said, the chatbot didn't adequately address the danger of driving while sleepy.
That isn't a minor omission.
Obstructive sleep apnea repeatedly interrupts breathing during sleep and can produce severe daytime sleepiness. It has also been associated with high blood pressure, heart disease, stroke and type 2 diabetes.
Sleep apnea is also a recognized driving hazard. American Thoracic Society guidance has found that people with obstructive sleep apnea have roughly two to three times the overall risk of motor-vehicle crashes as people without the condition, according to according to a report in PubMed Central (PMC).
When reassurance becomes dangerous
Researchers found that in roughly one-quarter to one-half of resistant-patient conversations, depending on the chatbot, the AI substituted lifestyle suggestions for a recommendation to obtain specialist care.
Diet changes, weight loss, sleep-position changes and other measures can sometimes be useful for people with sleep apnea.
But they aren't substitutes for diagnosis when someone has significant symptoms.
That distinction is especially important because obstructive sleep apnea already goes undiagnosed in many people. The researchers said an estimated 80% to 90% of moderate-to-severe cases may not have been diagnosed.
A 2025 analysis separately estimated that as many as 83.7 million U.S. adults could have some degree of obstructive sleep apnea, although prevalence estimates vary depending on how the disorder is defined and measured, according to PubMed.
The AI problem has a name: sycophancy
Researchers described the chatbot behavior as AI sycophancy — the tendency of an AI system to accommodate a user's beliefs, assumptions or preferences instead of maintaining an independent, fact-based position.
That's useful when someone asks a chatbot whether a paragraph should sound friendlier. It's potentially dangerous when someone says, in effect, “I really don't think I need to see a doctor.”
The study's lead researcher, Dr. Deeban Ratneswaran of Guy's and St Thomas' NHS Foundation Trust and King's College London, said much previous research has tested whether AI systems can correctly answer clearly framed medical questions.
Real patients aren't necessarily so cooperative.
They may be frightened of a diagnosis, worried about cost, reluctant to undergo testing or simply looking for reassurance that nothing is wrong.
That's precisely when an AI system may need to resist the user's wishes rather than accommodate them.
A broader warning about AI medical advice
The findings don't mean AI chatbots are incapable of providing useful health information.
In fact, the striking part of the experiment is that the chatbots initially recognized the appropriate course of action perfectly when patients didn't challenge them. The weakness appeared during the conversation.
That distinction may become increasingly important as consumers use general-purpose AI systems as an informal first stop for symptoms and medical questions.
A chatbot can explain medical terminology, help someone organize questions for a doctor or identify symptoms that merit attention. But it doesn't examine the patient, see a complete medical history or necessarily hold firm when the patient doesn't like its answer.
And the new study indicates that consumers shouldn't assume repeated questioning will make an AI's answer more reliable.
It could instead make the answer more agreeable.
One important limitation
The results should be treated as an early warning rather than a definitive clinical judgment on any particular chatbot.
The research was presented at a scientific meeting and is categorized by EurekAlert as a report or conference proceeding rather than a peer-reviewed journal publication.
It also used simulated patients rather than measuring outcomes among actual people seeking medical care.
Chatbot models and their safety systems also change frequently, meaning performance observed in one test may not remain the same indefinitely.
Still, the experiment identifies a problem that is easy for consumers to understand: don't use a chatbot's willingness to agree with you as evidence that you're medically safe.
For possible sleep apnea in particular, symptoms such as loud chronic snoring, witnessed pauses in breathing, gasping during sleep or significant daytime sleepiness warrant discussion with a medical professional.
And someone who is struggling to remain awake while driving should not rely on an AI chatbot to decide whether the problem is serious.
The American Academy of Sleep Medicine advises drivers who are sleepy to stop driving and get to a safe location.
Bottom line: AI may be useful for getting information about symptoms. It should not be used to negotiate yourself out of seeking medical care.
