Patient-facing chatbots and agents
Chatbots and voice agents that talk with patients directly, for symptom checking and triage, pre-visit intake and health-system phone calls.
These tools talk with patients directly: symptom checkers suggest possible conditions and how urgently to seek care, intake chatbots take a history before a visit, and voice agents call patients for health systems. FDA has said it intends to exercise enforcement discretion for symptom-checklist software that suggests possible conditions and when to see a clinician.
A 2022 systematic review of 10 studies found symptom checkers' primary-diagnosis accuracy ranged from 19% to 37.9% and triage accuracy from 48.8% to 90.1%. A five-year follow-up of 22 apps found median triage accuracy of 55.8% in 2020, close to 59.1% in 2015. In a randomized study of 1,298 members of the public, large language models alone identified the right conditions in 94.9% of scenarios, but participants using them identified relevant conditions in fewer than 34.5% of cases, no better than a control group. In a vignette test, OpenAI's consumer ChatGPT Health under-triaged 52% of gold-standard emergencies. Open questions include effects on where patients seek care and on their outcomes.
Ada
Ada asks users about their symptoms, then suggests possible conditions and how urgently to seek care; it is offered through a consumer app and through enterprise partners. According to the company, it is "an AI-powered symptom assessment optimized by clinicians."
Why this grade. Accuracy studies only: in the largest, a study of 437 ED patients in Germany who used Ada and another symptom checker in random order, Ada's first suggestion matched the discharge diagnosis for 14% and was identical or plausible for 58%; no study of its effect on care was found.
Caution. In the German emergency department study, Ada correctly triaged 34% of 437 patients, under-triaged 13% and over-triaged 53%, and it did not suggest a potentially life-threatening diagnosis for 13%; the authors called the missed diagnoses and inappropriate triage alarming.
The evidence
In the emergency department study at University Hospital Erlangen, six blinded experts compared the apps' suggestions with discharge diagnoses; Ada listed an identical diagnosis among its top five for 27% of patients and an identical or plausible one for 75%. In an accuracy study of 164 patients in a German rheumatology clinic, its top suggestion had 42.6% sensitivity and 63.6% specificity for inflammatory rheumatic disease. In a prospective accuracy study of 40 patients in an urban ED, its top five suggestions had 70% sensitivity for a final ED diagnosis, compared with 68.9% for physicians' top three, and at least two physicians rated 14% of its triage recommendations unsafe.
In an Ada-funded comparison on 200 primary care vignettes, in which all authors but one were current or former Ada employees or contractors or Ada equity holders, Ada's top three suggestions were correct for 70.5%, compared with an average of 82.1% for seven general practitioners and 23.5% to 43.0% for seven other apps.
- Used by
- Over 12 million people had used Ada's consumer app and over 50 million more had access through enterprise partnerships, the company said in December 2022.
- FDA
- No FDA clearance or FDA statement about Ada was found. FDA lists software that uses a checklist of signs and symptoms to suggest possible conditions and when to consult a health care professional among functions for which it intends to exercise enforcement discretion. In the European Union, Ada Assess was certified as a Class IIa medical device under the Medical Device Regulation in December 2022, the company said.
- Limits
- All the studies measured only how accurate the app is; none measured its effect on care or patient outcomes. Most were small or done in Germany.
Hippocratic AI care agents
Hippocratic AI's generative AI voice agents make and take patient calls for health systems: appointment access, post-discharge follow-up, chronic-care check-ins and care-gap outreach, with a human nurse one step away, according to the company.
Why this grade. All published evidence is funded and largely written by the company, including a peer-reviewed retrospective analysis of AI outreach calls to 1,878 WellSpan patients due for colorectal cancer screening, which found fecal immunochemical test kit opt-in of 18.2% among Spanish speakers versus 7.1% among English speakers.
The evidence
In the WellSpan Health analysis in September 2024, 18.2% of 517 Spanish-speaking patients called by the agent opted in to a test kit, versus 7.1% of 1,361 English-speaking patients (adjusted odds ratio 2.012); the study had no data on completed screening. Hippocratic AI funded it, and nine of its 12 authors were company employees.
In a company preprint, 6,234 U.S.-licensed clinicians reviewed more than 307,000 calls, and the rate of correct medical advice rose from about 80.0% before the company's Polaris models to 99.38% with Polaris 3.0. In a 2024 technical report posted as a preprint, the company said Polaris performed on par with human nurses in aggregate on simulated calls graded by more than 1,100 nurses and 130 physicians. A 2026 company preprint reported a clinical safety score of 99.9% across more than 10 million real patient calls.
- Used by
- WellSpan Health expanded its use platform-wide in July 2026, and its agent handled over 160,000 patient calls a month, HIT Consultant reported. University Hospitals in Cleveland announced a collaboration in September 2025. The company said it had completed more than 200 million real interactions.
- FDA
- No FDA clearance or authorization was found. The company said its agents do not diagnose or prescribe.
- Limits
- No independent study was found. The peer-reviewed study measured test kit opt-in at one health system over one month, without data on completed screening, and the safety figures above 99% come from company-run clinician grading on company-defined measures.
K Health
K Health's AI chat takes a patient's history before a virtual or in-person visit and suggests possible diagnoses and management to the treating clinician. According to the company, its dynamic intake and PatientGPT, a patient-facing chatbot, integrate directly into EHR systems for virtual and in-clinic care.
Why this grade. The evidence is retrospective: a K Health-funded cohort study of the AI in Cedars-Sinai Connect, a virtual urgent care clinic K Health lists as a partner, found initial AI and final physician recommendations agreed in 56.8% of 461 visits; eight of its 11 authors were from K Health, and no prospective or randomized study was found.
The evidence
In the retrospective cohort study of 461 adult visits to Cedars-Sinai Connect's virtual urgent care in June and July 2024, physician adjudicators rated the AI's initial recommendations optimal in 77.1% of visits and the treating physicians' final decisions optimal in 67.1%. The study could not determine whether physicians viewed the AI suggestions.
In an Ada-funded vignette comparison that tested K Health's 2020 consumer app, not the clinician-facing intake, its top three suggestions were correct for 36.0% of 200 vignettes, compared with an average of 82.1% for general practitioners.
- Used by
- Cedars-Sinai Connect, an AI-assisted virtual urgent care clinic, is among the partners K Health lists. Hackensack Meridian Health announced a 24/7 virtual primary care service with K Health in September 2024, Fierce Healthcare reported; K Health said 10 million people had interacted with its AI.
- FDA
- Neither an FDA clearance nor a company or FDA statement on its regulatory status was found. In the Cedars-Sinai Connect study, treating physicians made the final decisions.
- Limits
- The study was retrospective, vendor-funded and at one virtual clinic, and it measured agreement and rated quality rather than patient outcomes.
Grades: A, randomized evidence of benefit; B, evidence from clinical use; C, accuracy studies only; D, little or no independent evidence. How the grades work. Reviews of the published evidence, not medical advice or an endorsement of any product.