AI tools for clinical practice
31 tools, graded on the evidenceEvidence reviews of 31 AI tools physicians meet in practice, in nine categories, each graded on what published studies show.
Each review sets out what the tool does, who uses it, how the Food and Drug Administration (FDA) treats it and what published studies found, then gives an evidence grade from A to D and, where a credible independent study found poor real-world performance, a safety problem or harm, a caution flag. Every fact links to its source.
The grade measures the strength of the published evidence that a tool works in patient care, not the quality of the product. A low grade can mean the studies have not been done.
Ambient scribes and inbox drafts
Ambient scribes draft visit notes from the clinician-patient conversation, and inbox tools draft replies to patients' portal messages, for clinicians to review.
- AAbridge Abridge AI
Ambient AI scribe that listens to clinician-patient visits and drafts a structured note for the clinician to review.
- BEpic In Basket draft replies Epic Systems
Generative AI in Epic's In Basket that drafts replies to patient portal messages for clinicians to edit.
- BMicrosoft Dragon Copilot Microsoft
Microsoft's ambient AI scribe and dictation assistant, which drafts notes from recorded clinician-patient conversations.
- ANabla
Ambient AI assistant that captures the clinician-patient conversation and generates a structured clinical note.
Clinical questions and evidence search
AI search tools and chatbots that answer clinicians' medical questions in plain language, from the medical literature, an expert-written reference or a general-purpose model.
- CChatGPT OpenAI Caution
OpenAI's general-purpose chatbot, used by clinicians for clinical questions, with versions for individual clinicians and health systems.
- COpenEvidence
AI search tool that answers clinicians' questions with a summary of the medical literature and citations.
- CUpToDate Expert AI Wolters Kluwer Health
Generative AI chat feature in UpToDate that answers clinical questions from UpToDate's expert-authored, peer-reviewed content.
Patient-facing chatbots and agents
Chatbots and voice agents that talk with patients directly, for symptom checking and triage, pre-visit intake and health-system phone calls.
- CAda Ada Health Caution
AI symptom-assessment chatbot for consumers that suggests possible conditions and how urgently to seek care.
- DHippocratic AI care agents Hippocratic AI
Generative AI voice agents that make and take patient calls for health systems, from scheduling to follow-up.
- CK Health
AI chat intake that interviews patients before a visit and suggests possible diagnoses and management to the clinician.
Imaging: stroke and emergencies
Software that reads emergency imaging for suspected stroke, hemorrhage and other urgent findings and alerts radiologists and treatment teams.
- BAidoc Aidoc Medical Caution
Flags suspected hemorrhage, pulmonary embolism, fractures and other urgent CT findings and moves them up the radiologist's worklist.
- BRapidAI
Estimates infarct core and tissue at risk on perfusion imaging and flags suspected large vessel occlusions.
- AViz.ai grade for Viz LVO Caution
Detects suspected large vessel occlusion stroke on CT angiograms and alerts the stroke team; other modules cover hemorrhage.
Imaging: cancer and tuberculosis screening
Software that reads screening mammograms and chest X-rays, scores each exam for breast cancer, lung nodules or tuberculosis and marks suspicious areas.
- ALunit INSIGHT CXR Lunit grade for lung nodule detection
Reads chest X-rays and flags major thoracic abnormalities, including lung nodules, pneumothorax and pleural effusion.
- BLunit INSIGHT MMG Lunit
Scores 2D mammograms for likelihood of malignancy and highlights suspicious areas; can serve as an independent reader.
- BqXR Qure.ai
Reads chest X-rays for many findings and gives a tuberculosis score that selects people for confirmatory testing.
- ATranspara ScreenPoint Medical
Reads 2D mammograms and tomosynthesis, scores each exam's cancer risk and marks suspicious regions for the radiologist.
Cardiology
Software that reads ECGs, heart sounds and coronary CT scans to flag low ejection fraction, murmurs and coronary narrowings that limit blood flow.
- BAnumana ECG-AI LEF Anumana
Screens a routine 12-lead ECG for low ejection fraction in adults at risk for heart failure.
- BEko SENSORA Eko Health grade for low EF detection
AI on Eko's ECG-enabled digital stethoscopes that flags low ejection fraction, structural murmurs and atrial fibrillation.
- BHeartFlow FFRct and Plaque Analysis HeartFlow grade for FFRct
Simulates coronary blood flow and measures plaque from a coronary CT angiogram, showing which narrowings limit flow.
Sepsis and deterioration alerts
Software that uses EHR data to flag hospital patients at risk of sepsis or clinical deterioration.
- BBayesian Health TREWS Bayesian Health
Continuously monitors EHR data for hospital patients and flags those at risk of sepsis within 24 hours.
- BeCART AgileMD
An early warning score that estimates ward patients' risk of death or ICU transfer and guides care teams.
- BEpic Sepsis Model Epic Systems Caution
A sepsis risk score built into Epic's electronic health record that alerts clinicians above a hospital-set threshold.
- CSepsis ImmunoScore Prenosis
Combines 22 patient parameters, including sepsis biomarker levels, into a risk score for sepsis within 24 hours.
Point-of-care screening
AI devices that check for diabetic retinopathy at the point of care, or help primary care physicians decide whether to refer a skin lesion.
- CDermaSensor
Handheld AI device for primary care that scans skin lesions with light to help decide on dermatology referral.
- CEyeArt Eyenuk
Autonomous AI that grades retinal photographs for more-than-mild and vision-threatening diabetic retinopathy.
- ALumineticsCore Digital Diagnostics Caution
Autonomous AI that checks retinal photographs taken at the point of care for more-than-mild diabetic retinopathy.
Endoscopy and pathology
AI that flags possible polyps during colonoscopy, or suspicious tissue on digitized biopsy slides, for the physician to review.
- AEndoScreener Wision AI
Software that detects polyps in the colonoscopy video in real time for the endoscopist to confirm.
- AGI Genius Medtronic Caution
Computer-aided detection module that flags possible polyps on the colonoscopy video in real time.
- BIbex Galen Ibex Medical Analytics grade for Galen Breast
AI that reviews digitized pathology slides and flags tissue suspicious for cancer in breast and prostate biopsies.
- BPaige Prostate Detect Paige
Software that flags areas suspicious for cancer on digitized prostate biopsy slides for the pathologist to review.
How the grades work
- A
- Randomized evidence of benefit. Randomized controlled trials of the tool in routine care, published in peer-reviewed journals, met their primary outcome and found a benefit to patients or to care, such as more diagnoses made, fewer unnecessary procedures, better outcomes, or less clinician time or burden. Where such trials disagree, most of them did.
- B
- Evidence from clinical use. Studies have measured what happens when the tool is used in patients' care, such as its effect on diagnoses, decisions, workflow or outcomes, but randomized trials have not shown a consistent benefit, either because there are none or because their results are mixed or null.
- C
- Accuracy studies only. Studies measure only how accurate the tool is, in retrospective or prospective accuracy studies, test sets, simulated cases or reader studies, including the data behind an FDA clearance; none has measured its effect on care.
- D
- Little or no independent evidence. The published evidence comes only from the company, in its own studies, claims, white papers, case reports or press releases, with no FDA review of its data.
Caution Caution is a separate flag, shown with any grade, for a credible independent study that found poor real-world performance, a safety problem or harm, such as accuracy far below the company's own figures.
When a review covers several products or modules, the grade is for the one named beside it, and the review says how the others stand. A study counts for a tool if it tested that company's product, under its current or an earlier name or version, or an algorithm that the company, FDA or the algorithm's developers document as the one in the product or an earlier version of it; a predecessor developed separately does not count.
Grades now: A 8, B 14, C 8, D 1. Caution flags: 7.
About these reviews
FDA status is listed separately and is not evidence of benefit. A 510(k) clearance means FDA found a device substantially equivalent to one already legally marketed, and a De Novo authorization classifies a novel device that has no such predicate; neither is a finding that the device improves care.
The tools were chosen because they are widely used, widely marketed to physicians or widely studied, with three or four in each category. The reviews draw on peer-reviewed studies, preprints (labeled as such), FDA records, health technology assessments, news and institutional reports, and company pages; company statements are attributed and are not counted as evidence that a tool works. No company paid for, saw or approved its review before publication. The reviews are not medical advice or an endorsement of any product, and each shows the date it was last checked.
The evidence is current to Oct 4, 2026. Corrections and new studies can be sent through the contact form.
The Friday letter: the week that mattered and specific recommendations, free, in your inbox.