Ambient scribes and inbox drafts
Ambient scribes draft visit notes from the clinician-patient conversation, and inbox tools draft replies to patients' portal messages, for clinicians to review.
Ambient scribes listen to the visit and draft the note. Columbia Nursing, summarizing its researchers' commentary, said scribes are often classified as administrative tools rather than medical devices, allowing them to bypass FDA regulation. A Swedish health technology assessment found four randomized trials, all in the U.S., with 23 to 238 physicians each, and no data on total documentation time including editing, or on accuracy. Kaiser Permanente researchers estimated that scribes saved nearly 16,000 hours of documentation time over more than 2.5 million uses. In a MedStar Health simulation, 31 of 44 draft notes contained errors.
Inbox tools draft replies to messages from patients, which rose to 157% of the prepandemic average for U.S. ambulatory clinicians during the pandemic. Of 23 studies in a 2025 systematic review of AI-drafted replies, 16 were simulations; all seven in live use tested GPT-4 tools built into Epic. In a Mass General Brigham simulation, 7.1% of GPT-4 drafts could pose a risk to the patient if sent unedited, and 0.6% a risk of death.
Abridge
Abridge listens in the background of a clinician-patient visit and drafts a structured, specialty-specific note inside the EHR for the clinician to review and sign; the company says the note is ready for review the moment the visit ends.
Why this grade. A stepped-wedge randomized trial of 66 practitioners at UW Health (Afshar et al., NEJM AI, 2025) found that the scribe, which a University of Wisconsin release identifies as Abridge, reduced work exhaustion, one of two primary outcomes, and time on notes by 0.36 hours a day; professional fulfillment did not improve significantly.
The evidence
A stepped-wedge randomized trial of 66 physicians and advanced practice providers at UW Health clinics in two states found that ambient AI reduced work exhaustion and interpersonal disengagement by 0.44 points (P<0.001) and time on notes by 0.36 hours a day; professional fulfillment did not improve significantly. The paper's abstract does not name the product; a University of Wisconsin release identifies it as Abridge.
At Sutter Health, a before-and-after pilot of 100 clinicians used a tool that Sutter identifies as Abridge: time in notes per appointment fell from 6.2 to 5.3 minutes, and the share reporting burnout fell from 42.1% to 35.1%, a change that was not significant. In interviews after the pilot, clinicians reported missed or inaccurate information that required them to review the transcript or audio and edit the note; Sutter Health, which funded that study, is a strategic investor in Abridge.
- Used by
- More than 300 U.S. health systems use it, Abridge said in August 2026; the company says Sutter Health has more than 3,000 users. Kaiser Permanente announced a rollout at 40 hospitals and more than 600 medical offices in August 2024.
- FDA
- No FDA clearance was found. AI scribes are often classified as administrative tools rather than medical devices, allowing them to bypass FDA regulation, according to a Columbia Nursing summary of a commentary by its researchers in npj Digital Medicine.
- Limits
- The trial was unblinded, at one academic health system, with 66 practitioners and self-reported primary outcomes. The Sutter pilot had no concurrent control group and was run at a health system that invests in Abridge.
Epic In Basket draft replies
Generative AI built into Epic's In Basket drafts potential replies to patients' portal (MyChart) messages. The clinician can start from the draft and edit it, or write a reply from scratch, before sending.
Why this grade. The one randomized study, of 52 primary care physicians at UC San Diego Health (Tai-Seale et al., JAMA Network Open, 2024), found drafts were associated with significantly longer read times and no change in reply time; prospective studies at Stanford Health Care and a Dutch academic hospital found drafts were used without a change in response time.
The evidence
In a randomized waiting-list study of 52 primary care physicians at UC San Diego Health, draft access was associated with significantly longer read time, no change in reply time and significantly longer replies. In a five-week pilot of 162 clinicians at Stanford Health Care, mean draft use was 20%; task load and work exhaustion scores fell, and reply, write and read times did not change. In a 16-week study of 100 clinicians at a Dutch academic hospital, drafts were used in 58% of replies, and total response time did not differ between blank replies and replies built from drafts (157 seconds versus 153 seconds).
An audit-log analysis at NYU Langone Health, covering 55,767 patient messages, found drafts were used in 19.4% of 5,935 eligible instances, and drafts cut turnaround time by 6.76%. In a blinded comparison at NYU Langone, 16 primary care physicians found no statistical difference in accuracy, completeness and relevance between AI drafts and clinicians' replies; the drafts were 38% longer.
- Used by
- Epic's chief executive, Judy Faulkner, said in August 2024 that 150 health systems and medical groups were using it, generating 1 million drafts a month. Epic says Mayo Clinic uses it and that initial pilots saved nurses around 30 seconds per message.
- FDA
- No FDA clearance for it was found in FDA's 510(k) database, and no source found states how FDA classifies AI-drafted replies to patient messages.
- Limits
- Each study was at one health system, burden outcomes were self-reported, the one randomized study had 52 physicians, and no study has measured how often draft errors reach patients in live use. In a Pennsylvania pilot, 40% of replies showed little to no editing of the draft, and 15 free-text responses (34.9%) said drafts occasionally contained incorrect or inappropriate content.
Microsoft Dragon Copilot
Dragon Copilot, announced by Microsoft in March 2025, combines the ambient listening of Dragon Ambient eXperience (DAX), which drafts a specialty-specific note from the recorded clinician-patient conversation, with Dragon Medical One voice dictation.
Why this grade. Randomized trials are mixed: the largest, a three-arm trial of 238 physicians at UCLA Health, found no significant change with DAX in time-in-note, the trial's primary outcome, while smaller trials at Providence and in pediatric subspecialty clinics found less burden or burnout and a 45-clinician pilot trial found no significant change.
The evidence
In a three-arm randomized trial of 238 physicians at UCLA Health, time-in-note fell 1.7% with DAX versus usual care, a change that was not significant (P=0.66); burnout, task load and work exhaustion improved with either scribe, in secondary outcomes that need confirmation. DAX was used in 33.5% of visits, the trial's preprint reported. Smaller randomized trials of 24 providers at Providence and 23 pediatric subspecialists found lower burnout with DAX; Providence providers spent 2.5 hours less a week on off-hours documentation, while the pediatric trial found no change in charting time. A Swedish health technology assessment rated both trials at unacceptably high risk of bias in randomization. In the STREAMLINE pilot trial of 45 clinicians in community primary care, even among frequent users a 1.4-minute drop in time on notes per problem-focused visit was not significant.
A matched cohort study at Intermountain Health, using data from March to September 2022, compared 99 DAX users with 76 controls and found positive trends in engagement and no practical change in productivity.
- Used by
- DAX assisted more than 3 million patient conversations at 600 organizations in one month, Microsoft said in March 2025. Intermountain Health has more than 2,500 active clinician users, Microsoft says. Providence deployed it enterprise-wide for ambulatory clinicians, a 2025 Peterson Health Technology Institute report said.
- FDA
- No FDA clearance was found. AI scribes are often classified as administrative tools rather than medical devices, allowing them to bypass FDA regulation, according to a Columbia Nursing summary of a commentary by its researchers in npj Digital Medicine.
- Limits
- Nuance supplied the licenses for the Providence trial but had no role in its design or analysis, most reported benefits are self-reported, and the Swedish assessment found no randomized data on total documentation time including editing, or on accuracy.
Nabla
Nabla captures the patient-clinician conversation and generates structured clinical notes automatically, according to the company, which says it works inside Epic and other EHRs and offers speech recognition and coding suggestions.
Why this grade. In a three-arm randomized trial of 238 physicians at UCLA Health (Lukac et al., NEJM AI, 2025), time-in-note, the primary outcome, fell 9.5% with Nabla versus usual care (P=0.02); the trial's preprint reported an estimated drop of 41 seconds a note in the Nabla arm and 18 seconds in the control arm.
The evidence
In a three-arm randomized trial of 238 physicians in 14 specialties at UCLA Health, time-in-note fell 9.5% with Nabla versus usual care (P=0.02); the trial's preprint reported Nabla was used in 29.5% of visits. Burnout, task load and work exhaustion improved with either scribe, in secondary outcomes that need confirmation. The trial noted occasional clinically significant inaccuracies and reported one mild (grade 1) adverse event.
In a five-week pilot of 38 physicians and advanced practice providers at University of Iowa Health Care, with no control group, the median burnout score improved from 4.16 to 3.16 (P=0.005) and the burnout rate fell from 69% to 43%.
- Used by
- Nabla is deployed in more than 130 health organizations, the company says. University of Iowa Health Care has used it system-wide since September 2024.
- FDA
- No FDA clearance was found. AI scribes are often classified as administrative tools rather than medical devices, allowing them to bypass FDA regulation, according to a Columbia Nursing summary of a commentary by its researchers in npj Digital Medicine.
- Limits
- The benefit rests on one unblinded trial at one academic health system, and the Iowa pilot had no control group; a Swedish health technology assessment found no randomized data on total documentation time including editing, or on accuracy.
Grades: A, randomized evidence of benefit; B, evidence from clinical use; C, accuracy studies only; D, little or no independent evidence. How the grades work. Reviews of the published evidence, not medical advice or an endorsement of any product.