Clinical questions and evidence search
AI search tools and chatbots that answer clinicians' medical questions in plain language, from the medical literature, an expert-written reference or a general-purpose model.
These tools answer clinicians' medical questions in conversational form. In a 2026 American Medical Association survey of nearly 1,700 physicians, 81% reported using AI professionally, more than double the 2023 rate. FDA's guidance on clinical decision support software describes the types of decision-support functions excluded from the definition of a device; no FDA clearance was found for any of the three tools.
A systematic review of 519 studies of large language models in health care found 5% evaluated them on real patient care data. A 2026 review of 4,609 such studies in clinical medicine found 1,048 that used real-world patient data, 19 of them prospective randomized trials. In NOHARM, a preprint that tested 20 large language models and four retrieval-based clinical tools on a 1,100-task benchmark of primary care-to-specialist consultation cases, applying the recommendations directly carried potential for severe harm in up to 24.6% of cases; omissions accounted for more than 80% of severe errors. Open questions include whether the answers change clinicians' decisions and patient outcomes in routine care.
ChatGPT
ChatGPT is OpenAI's general-purpose chatbot; the review covers it and the GPT models behind it, such as GPT-4 and GPT-4o, but not other companies' tools built on them. ChatGPT for Clinicians, a version released in April 2026, is "designed to support clinical tasks like documentation and medical research," according to OpenAI.
Why this grade. The randomized trials of GPT-4 used simulated cases: physicians using GPT-4 through ChatGPT Plus scored 6.5 percentage points higher on management reasoning in a trial of 92 physicians, and GPT-4 did not significantly improve diagnostic reasoning in a trial of 50 physicians; the trials in live care tested AI Consult, a different GPT-4o-based tool.
Caution. In a simulation study of 300 physician-validated vignettes, each containing one fabricated detail, GPT-4o, the top performer among six models, repeated or elaborated on the false detail in 53% of cases, and in 23% with a mitigation prompt. In a randomized vignette trial of 44 physicians in Pakistan, reported in a preprint, those shown deliberately flawed ChatGPT-4o recommendations had mean diagnostic accuracy of 73.3%, versus 84.9% with unmodified ones.
The evidence
In a randomized trial of 92 physicians working through five simulated cases, those given GPT-4 through the ChatGPT Plus interface scored 6.5 percentage points higher on management reasoning than those using conventional resources (95% confidence interval 2.7 to 10.2). In a randomized trial of 50 physicians on simulated diagnostic cases, median scores were 76% with GPT-4 and 74% without, not a significant difference; GPT-4 alone scored 16 percentage points higher than physicians using conventional resources.
The trials in live care tested AI Consult, a GPT-4o-based safety net that Penda Health built into its records system in Kenya, not ChatGPT. A cluster-randomized trial of 103 clinical officers at 16 Penda clinics, with 9,691 patients, found no significant reduction in treatment failure within 14 days, its primary outcome (adjusted odds ratio 0.77, 95% confidence interval 0.55 to 1.08). An OpenAI-funded quality-improvement study of 39,849 visits, in which half the clinicians in each clinic were randomly given access, reported 16% fewer diagnostic errors and 13% fewer treatment errors; it is a preprint.
- Used by
- Millions of clinicians worldwide used ChatGPT in clinical care every week, OpenAI said in April 2026. ChatGPT for Healthcare, OpenAI's version for health systems, launched Jan. 8, 2026, and was rolling out to institutions including AdventHealth, Cedars-Sinai Medical Center and HCA Healthcare, the company said.
- FDA
- No FDA clearance or authorization was found; ChatGPT is general-purpose software. OpenAI said ChatGPT for Clinicians "is designed to support clinicians with information, not replace their judgment or expertise."
- Limits
- No study was found of patient outcomes when clinicians use ChatGPT, ChatGPT for Clinicians or ChatGPT for Healthcare in practice; the vignette trials measured reasoning scores. OpenAI funded the study that reported fewer errors with AI Consult and took part in its analysis and reporting.
- Cost
- ChatGPT for Clinicians is free for verified physicians, nurse practitioners, physician assistants and pharmacists, starting in the U.S., OpenAI said. No price is published for ChatGPT for Healthcare.
OpenEvidence
OpenEvidence is a clinical question-answering platform: clinicians ask a question and receive a summarized answer with citations to the medical literature. According to the company, it is an AI system built to aggregate, synthesize and visualize clinically relevant evidence in understandable, accessible formats.
Why this grade. A systematic review of 11 early evaluations of OpenEvidence (Artsi et al., npj Digital Medicine 2026), all testing the quality of its answers, found generally clinically relevant, evidence-supported responses and called for prospective real-world studies; no prospective study of the tool in clinical use was found.
The evidence
The systematic review found OpenEvidence had lower rates of fabricated citations than general-purpose large language models and was strongest on guideline-based questions, variable in complex cases. In an independent NYU Langone Health benchmark, three frontier general-purpose models outperformed OpenEvidence and UpToDate Expert AI on exam questions, HealthBench items and 100 real physician queries; on the real queries, the clinical tools performed comparably to Google Search's AI Overview. In a preprint study, 149 physicians rated OpenEvidence above Claude Opus 4.8, Gemini 3.1 Pro and GPT-5.5 on all five measures; OpenEvidence supplied the queries, ran the data collection and paid the raters, and the paper said no author was affiliated with the company.
In a Mayo Clinic review of five primary care cases, four physicians rated its answers 3.75 for relevance and 1.95 for impact on decisions on a 0 to 4 scale; the answers mostly reinforced existing plans. An audit of 4,979 cited references confirmed none as fabricated. On 15 questions about tricuspid valve therapies, two cardiac imaging specialists rated 53.3% of its answers inaccurate, compared with 13.3% for ChatGPT-4o.
- Used by
- More than 40% of U.S. physicians used it daily, on average, and it supported about 18 million clinical consultations in December 2025, the company said in January 2026. Its use spanned more than 10,000 hospitals and medical centers, according to the company.
- FDA
- No FDA clearance or authorization was found. According to the company's terms, OpenEvidence "does not provide medical advice, diagnosis or treatment." FDA's clinical decision support guidance describes the types of decision-support software functions excluded from the definition of a device; no FDA determination about OpenEvidence was found.
- Limits
- The evaluations were small and heterogeneous, none measured effects on care or patient outcomes, and the platform changes continuously. Head-to-head results conflict: an independent benchmark ranked it below general-purpose models, while a preprint for which OpenEvidence ran the data collection and paid the raters ranked it first.
- Cost
- Free to verified U.S. clinicians and paid for by pharmaceutical and medical device advertising, according to an analyst profile by Sacra. The company said OpenEvidence is "free to doctors and ad-supported."
UpToDate Expert AI
UpToDate Expert AI is a chatbot-like feature of UpToDate that gives generative AI answers to clinical questions based on UpToDate's expert-authored and peer-reviewed content, according to Wolters Kluwer. The company said it would be available to select UpToDate Enterprise Edition customers beginning in the fourth quarter of 2025.
Why this grade. Independent evidence is limited to a benchmark study by NYU Langone Health researchers (Vishwanath et al., Nature Medicine 2026), in which three frontier general-purpose models outperformed it on exam questions, HealthBench items and 100 real physician queries; the company's 99.9% figure comes from its own white paper.
The evidence
In the NYU Langone Health benchmark, GPT-5.2, Gemini 3.1 Pro and Claude Opus 4.6 outperformed UpToDate Expert AI and OpenEvidence on 500 MedQA exam questions, 500 HealthBench items and 100 real physician queries rated by 12 U.S. clinicians. On the real queries, the two clinical tools performed comparably to the AI Overview that Google Search enables automatically. In a white paper, Wolters Kluwer reported that the tool gave clinically aligned information for 99.9% of assessed criteria in internal testing on 1,669 clinical queries; the document is not peer reviewed. No prospective study of the tool in clinical use was found.
- Used by
- About 2,500 hospitals and health systems, representing over 90% of U.S. UpToDate Enterprise Edition customers, had chosen it, Wolters Kluwer said in August 2026; it named Sharp HealthCare, Duke University Health System, University of Iowa Health Care and Intermountain Health among them.
- FDA
- No FDA clearance or authorization was found, and Wolters Kluwer's launch and adoption announcements make no regulatory statement. FDA's clinical decision support guidance describes the types of decision-support software functions excluded from the definition of a device; no FDA determination about UpToDate Expert AI was found.
- Limits
- No study was found that measured its effect on clinicians' decisions or patient outcomes, and the one independent evaluation is a benchmark rather than a study in clinical use. The company's 99.9% figure comes from internal testing described in a white paper that is not peer reviewed.
Grades: A, randomized evidence of benefit; B, evidence from clinical use; C, accuracy studies only; D, little or no independent evidence. How the grades work. Reviews of the published evidence, not medical advice or an endorsement of any product.