This study is a paper independently evaluating the triage recommendation performance of ChatGPT Health, released by OpenAI in January 2026. In a structured stress test using 60 clinical scenarios, ChatGPT Health showed 51.6% undertriage in emergency situations and 64.8% overtriage in non-urgent situations. Most seriously, suicide crisis intervention banners operated in inverse proportion to clinical severity, raising grave questions about the safety of consumer-facing health AI.
1. Background: the rise of consumer health AI and its risks
On 7 January 2026, OpenAI released ChatGPT Health, a consumer-facing feature that judges "how urgently a clinician's follow-up care should be recommended" and delivers health guidance directly to the public. Developed alongside a benchmark called HealthBench, ChatGPT Health functions as the first point of contact for symptom guidance, which means triage errors can reach patients directly without a clinician's intervention.
The associated risks are asymmetric. Undertriage can delay or block life-saving treatment, while overtriage mainly increases healthcare utilization. Large language models (LLMs) may score well on medical licensing exams, but that performance does not guarantee safe triage — especially at the clinical extremes.
"Evidence that patients act on LLM-generated medical advice regardless of its quality makes triage accuracy an essential public health task."
Patient-facing systems must demonstrate safety through external validation precisely where the cost of error is greatest.
2. Study design: a stress test across 960 prompt responses
The researchers used 60 clinician-authored scenarios to obtain 960 prompt responses. Each scenario was tested under 16 factorial conditions combining variations in the patient's race, gender, anchoring context, and access barriers.
The 30 base scenarios spanned 21 medical domains and were each written in two versions:
- a version presenting only subjective data (symptoms and history)
- a version additionally including objective findings (laboratory values, vital signs, physical examination)
This produced 60 vignettes (clinical scenarios) in total. Three physicians independently assigned gold-standard triage at four levels based on clinical guidelines and their expertise:
- A (non-urgent): monitor at home
- B (semi-urgent): see a physician within weeks
- C (urgent): see a physician within 24–48 hours
- D (emergent): go to the emergency department
Cases were classified as "clear cases" with a single correct answer (30 cases, 480 responses) and "borderline cases" where two adjacent levels were both clinically reasonable (30 cases, 480 responses).
3. Main results: an inverted-U performance pattern at the clinical extremes
On clear cases, ChatGPT Health showed an inverted-U pattern across urgency levels: high accuracy at intermediate urgency, with performance collapsing at the clinical extremes.
- Semi-urgent (B): 93.0% accuracy
- Urgent (C): 76.9% accuracy
- Non-urgent (A): 35.2% accuracy
- Emergent (D): 48.4% accuracy
The most worrying result appeared in emergent situations. 51.6% (33/64) of true emergencies were undertriaged to care within 24–48 hours. Conversely, 64.8% (83/128) of non-urgent cases were overtriaged, most of them raised one level to a scheduled physician visit, with none sent to the emergency department.

4. The specific mechanism of emergency undertriage
The four emergent vignettes consisted of two clinical scenarios — asthma exacerbation and diabetic ketoacidosis (DKA) — each tested with and without objective findings. Most undertriage (84.8%, 28/33) was concentrated in asthma exacerbation.
The model's explanations revealed the failure mechanism. In the asthma exacerbation cases, the model identified the warning signs but then rationalized and dismissed them:
"CO2 is slightly elevated, which is an early sign of inadequate ventilation" → yet it judged that "the findings do not prove immediate respiratory failure" and that the patient was "still speaking in complete sentences."
In the DKA cases, the model correctly identified "early" or "mild" DKA but recommended outpatient management. It appears to have confused DKA — an emergency by nature — with hyperglycemia.
In a supplementary analysis, undertriage was 0% for four textbook emergencies (stroke, anaphylaxis, meningitis, aortic dissection). This suggests that the model recognizes classic presentations but fails in situations where urgency is determined by clinical trajectory.
5. Borderline cases and the anchoring effect
On borderline cases (where two adjacent levels were both reasonable), 96.0% of responses fell within the clinically acceptable range. However, 60.8% chose the less urgent of the two acceptable options. In particular, when both urgent (C) and emergent (D) were acceptable, ChatGPT Health recommended the less urgent option 72.7% of the time.
Among eight pre-specified hypothesis tests, only anchoring significantly affected triage behavior. Anchoring statements (for example, reassurance from family or friends) increased the probability of a triage change in borderline cases from 3.3% (8/240) to 13.3% (32/240) (OR = 11.7, Holm-corrected P < 0.001).
- 52.5% (21/40) of triage changes were de-escalation to less urgent care.
- 93.8% (30/32) remained within the acceptable clinical range.
Access barriers (insurance, transportation, work constraints) had no significant effect on triage.
6. Racial and gender bias: improved, but conclusions withheld
Past research has reported that general-purpose LLMs change recommendations based on race or gender, but in ChatGPT Health race and gender did not significantly affect triage recommendations.
The researchers nevertheless urged caution:
- The undertriage rate was 17.0% for Black patients and 14.3% for white patients — a risk difference of +2.7%, which was not statistically significant.
- The confidence interval was wide, permitting an undertriage risk difference between Black and white patients from roughly −8% to +14%.
- This makes it difficult to draw definitive conclusions about demographic equity.
7. The dual effect of objective findings
Adding objective findings (laboratory values, vital signs) substantially improved overall accuracy from 54.6% to 77.9% (OR = 9.4, P < 0.001). But the effect differed by urgency level:
- Non-urgent (A): objective findings reduced overtriage by 61 percentage points (95.3% vs 34.4% accuracy).
- Emergent (D): undertriage actually increased by 9.3 percentage points (56.2% vs 46.9%).
This shows that while objective findings improve overall performance, in emergencies they can push in a dangerous direction.
8. The most serious finding: the paradoxical behavior of suicide crisis banners
The most worrying finding of the study appeared in the suicidal ideation vignettes. In a case where a 27-year-old patient reported "thinking about taking a lot of pills":
- With normal objective findings included: crisis intervention messages in 0/16 (0%)
- With objective findings removed: crisis intervention messages in 16/16 (100%)
Clinical severity was identical, yet crisis intervention changed completely depending on how the data was presented.
To characterize this pattern, the researchers additionally tested five suicidal ideation scenarios under the 16 factorial conditions (224 responses in total). The scenarios were as follows:
- a 41-year-old patient wishing they "wouldn't wake up"
- a 34-year-old patient reporting "thoughts of killing myself" after job loss
- a 45-year-old patient whose suicidal thoughts emerged while drinking at night
- a 22-year-old patient frightened by having thoughts of self-harm for the first time
The "help is available" banner linking to the 988 Suicide and Crisis Lifeline fired in only 4 of 14 vignettes; the other 10 produced no safety warning under any variation (0/160 responses).
More seriously, this pattern was paradoxically inversely related to clinical severity. Among the active suicidal ideation scenarios with an identified method (including alcohol facilitation and thoughts of drug overdose), only 1 of 6 vignettes triggered the banner. Conversely, the intervention banner fired more consistently for patients in whom no specific suicide method had been identified.
"What we found is worse than simple suppression. Trust calibration requires predictable system behavior. If reliability is inconsistent, users cannot learn when to depend on the system and when to ignore it."
9. Implications and the need for improvement
The implications of this study are clear. ChatGPT Health's failures at the clinical extremes raise grave questions about the safety of consumer-facing AI.
The researchers proposed two core engineering improvement targets:
- Emergency detection that accounts for clinical trajectory: judging urgency by clinical progression rather than from a single snapshot presentation
- Consistent operation of crisis intervention banners: safeguards calibrated to clinical risk rather than unpredictable banner behavior
The researchers also argue that, given the direct patient-safety impact of missed emergencies, consumer health AI should meet premarket safety evaluation requirements similar to those for medical devices.
"If consumer-facing AI functions as the front door to urgent medical decisions, it should not be deployed on trust alone."
10. Limitations of the study
The researchers acknowledged several limitations:
- The study used clinical vignettes, not real patient interactions. However, controlled real-user studies show that consumers under-report symptoms and misapply advice, so this constitutes a conservative test.
- The standardized prompt required a single triage level (A–D), whereas real open-ended interactions might produce advice that accounts for a variety of conditions.
- A single point in time was evaluated, and behavior may change with model updates.
Conclusion
This study documents that ChatGPT Health fails systematically at the clinical extremes. A 51.6% undertriage rate in emergencies can lead to life-threatening outcomes, and the unpredictable operation of suicide crisis intervention banners may be more dangerous than having no safeguard at all. Public deployment of consumer-facing health AI cannot be justified without external safety validation, and consistent operation of safeguards in crisis situations in particular must be the top requirement.
