Evidence

13 items from the daily posts

What the studies and the professional societies say, and what nobody has shown yet. Each item links to the day it ran and to its original source.

  • Evidence Sept 22, 2026
    Xu and colleagues at Zhongshan Hospital in Shanghai reported in npj Digital Medicine on FARUSS, a six-axis robotic arm that positions the probe and scans the thyroid on its own using deep learning, feeding an AI platform that detects nodules and assigns TI-RADS categories. The team tested it prospectively against conventional scans in 262 participants at three Chinese institutions from March 2024 to August 2025: intraclass correlations for thyroid measurements were 0.76 to 0.80, nodule identification matched 85.6% of the time, and TI-RADS agreement after a sonologist revised the AI's call reached kappa 0.75 to 0.88. The robot was slower to scan, 202 seconds against 147, and the AI read faster than a human, 183 seconds against 254. In the authors' triage scheme, the system avoided 74.8% of on-site examinations that turned out not to be needed and missed 19.2% of the fine-needle aspiration recommendations a human would have made. Three authors work for the technology companies involved, and the study was funded by the Chinese government.
  • Evidence Sept 22, 2026
    Zhou, Guo, Wang and colleagues trained a two-stage model, one stage to locate the esophagus and one to segment lesions and score malignancy, on 6,813 patients from two centers, then validated it on 80,612 patients across 12 centers in China, the Czech Republic and Australia. On an external test set of 11,466 patients from eight centers, the model found esophageal cancer with 90.0% sensitivity and 98.5% specificity and high-grade precancer with 52.5% sensitivity; on 1,607 low-dose chest CTs it ran at 88.4% sensitivity and 99.0% specificity. In a reader study it outperformed all 17 radiologists, exceeding their average by 19.1 points of sensitivity and 18.4 points of specificity, and as an assistant it raised the radiologists' cancer sensitivity from 71.9% to 85.7%. Run prospectively on hospital scans from January to April 2025, followed through July 2026, and on 10,959 consecutive asymptomatic lung screening participants, the model's flags had a positive predictive value of 42.2%, with specificity of 99.94% in the screening cohort. The authors listed the study's limits: nearly all patients were Chinese, a population in which squamous cell cancer dominates, with only preliminary validation abroad; performance was worse in women because the training set skewed male; follow-up was less than two years with imperfect compliance; and the reader study lacked a full crossover design. In the U.S. the common tumor is adenocarcinoma of the distal esophagus, which the model has not been shown to detect. The paper appeared one day after Alibaba's abdominal CT model was released free of charge.
  • Evidence Sept 16, 2026
    The 49% figure comes from a special analysis of the 2026 Edelman Trust Barometer written with the Yale School of Public Health and released this month, which Becker's Hospital Review reported on Monday. Edelman surveyed 12,998 people in 13 countries, about 1,000 per country including 995 Americans, between Feb. 28 and March 11. Asked which tasks a person with no medical training but skilled with AI could do as well as a trained professional, 26% chose deciding whether someone needs care, 19% chose determining treatment or medication, 19% chose performing basic procedures and 16% chose diagnosing illness. Agreement ran 59% among ages 18 to 34, 52% among ages 35 to 54 and 35% among those over 55, and 56% among the university-educated against 42% among those without a university education. The same analysis found that only 51% of people who receive conflicting recommendations always follow their doctor's. Edelman is a public relations firm and the barometer is its product; the analysis was written with Yale and its method is published. The question measured what respondents believe an AI-skilled layperson could do, not what such a person can do.
  • Evidence Sept 18, 2026
    The model was trained on more than 400,000 contrast-enhanced abdominal CT exams paired with their reports, about 15 million anatomy-tagged image-text pairs, according to the GitHub page, which links to the Science paper. The South China Morning Post reported Thursday that the model was tested on nearly 40,000 real-world exams and recorded a mean AUC of 0.913 across 146 findings in 18 organs, cancers included. The test set came from eight outside centers, the model outperformed 23 of 26 radiologists, and radiologists using it as an assistant gained about 10 points of sensitivity and read more than 30% faster, TechTimes reported; the code is released under the Apache 2.0 license and the weights are licensed for research, not commercial use. The Science paper is behind a paywall, and the figures here come from the company's page and the press reports. The model was validated on Chinese patients only and on contrast-enhanced abdominal CT only, with no prospective trial and no regulatory clearance in any country. Under a final order published Thursday in the Federal Register, every radiology detection, diagnosis and triage algorithm in the U.S. still requires its own 510(k) clearance from the FDA, open weights or not.
  • Evidence Sept 19, 2026
    Emergency physicians and a bioinformatician at UT Southwestern built pairs of patients from two academic emergency departments, matched on Emergency Severity Index, in which one patient deteriorated within six hours and the other did not, then asked Gemma, Qwen and DeepSeek which to see first, before and after inserting stigmatizing or neutral versions of 18 attributes. Stigmatizing language about frequent emergency department use pushed the deteriorating patient down the list in every model and dataset, by as much as 9.4%, and neutral phrasing scored 1.4 to 7.1 points higher, according to the paper in npj Digital Medicine. Psychiatric history moved rankings by up to 9.5% but reached significance in only one model, and race, language and insurance showed no consistent harm. The authors concluded that the wording of the input is a safety-critical design choice. Thursday's post covered a model that scored identical presentations as less severe when the patient was female.
  • Evidence Sept 21, 2026
    Radiologists at St. Jude reviewed every chest radiograph foundation model published through March 2025, 41 in all, and scored each paper against CLAIM, the reporting checklist for medical imaging AI. The brief communication in npj Digital Medicine found transformer architectures and language-model components becoming dominant, model sharing still uncommon and one variable that tracked with quality: publications with a physician among the authors adhered to CLAIM significantly better and discussed fairness more often. The paper is short, and its authors report the link as a correlation only.
  • Evidence Sept 17, 2026
    The study, published in NEJM AI, examined a year of Epic-integrated draft replies to patient portal messages across specialties at UC San Diego Health. Researchers sorted every physician edit into a 15-category taxonomy built with language models and expert review and used response time as the measure of workload. Scheduling changes were the most common edit, appearing in 38.5% of edited messages, followed by lifestyle and non-drug advice in 18.2% and added empathy in 16.1%; stopping or tapering a medication was the rarest edit at 2.4%. Clinical edits added the most time per message: interpreting a radiology result added 70.1% to the time spent on a message, clarifying or ruling out a diagnosis added 63.9% and interpreting laboratory results added 60.8%. The authors concluded that the burden is not uniform, with judgment edits costing the most per message and high-volume administrative edits costing the most across the system, according to reports in Becker's Physician Leadership and Healthcare Innovation; the paper is behind NEJM AI's paywall, and the figures come from those two reports.
  • Evidence Sept 11, 2026
    Kohn and colleagues ran Epic's one-year mortality model, which the paper describes as "perhaps the most widely available commercial mortality risk model," on 154,063 encounters at Trinity Health and 133,043 at Kaiser Permanente Southern California from 2022 through 2023. Discrimination was moderate to high, with a C statistic of 0.76 at Trinity and 0.81 at Kaiser, but calibration was poor in both systems, and the scaled Brier score at Trinity was essentially zero, meaning the probabilities the model reported there were about as useful as the base rate. Performance was weakest in patients 75 and older, where the C statistic fell to 0.69 and 0.73, and it degraded in liver disease, kidney disease and heart failure. Dr. James Deardorff, a UCSF geriatrician, wrote in STAT on Friday that the same prediction is harmless when it prompts a goals-of-care conversation and dangerous when it feeds a transplant or resource decision. Many hospitals have the score switched on.
  • Evidence Sept 16, 2026
    Guerra-Adames and colleagues at Bordeaux University Hospital trained a Mistral model to reproduce nurses' triage scores, then fed it pairs of cases that differed only in the patient's sex, they reported in npj Digital Medicine. Female versions came out 1.1% lower in predicted severity in Bordeaux and 2.2% lower on the U.S. MIMIC-IV records; retraining on sex-neutralized text made the gap disappear, and the pattern shifted with the sex of the triage nurse. The authors described the result as a probe of what the charts record rather than proof of bedside undertriage, and as a hypothesis generator rather than a verdict.
  • Evidence Sept 17, 2026
    A group from the Harvard T.H. Chan School of Public Health and the University of Pennsylvania published the rapid review in npj Digital Medicine on Sept. 17, covering 229 randomized trials from 29 systematic reviews in nutrition, maternal health, mental health and sleep from 2013 to 2025. Developers took part in 73% of the trials. Developer-involved trials had 2.47 times the odds of being preregistered and, when weighted by sample size, higher odds of reporting a statistically significant result, an odds ratio of 1.23 with a confidence interval of 1.16 to 1.31. The authors called for independent efficacy trials, more transparent reporting and regulatory oversight. The domains studied were behavioral rather than diagnostic; the review did not cover ambient scribes or diagnostic agents.
  • Evidence Sept 15, 2026
    Kather's group tested a fully autonomous agent that interviews the patient, examines, orders tests and diagnoses on 551 MIMIC-IV cases across seven acute conditions (appendicitis, cholecystitis, diverticulitis, pancreatitis, pneumonia, pulmonary embolism and urinary infection), 2,400 abdominal cases and 990 published multispecialty cases. Open-weight models hosted on premises (Qwen, GLM and GPT-OSS), which keep patient data inside the hospital, reached 90.0% on the seven-disease task and 83.8% on the four-disease task, near a cloud GPT-5.2 baseline, the authors said. Agreement across five repeated runs separated right from wrong answers with an AUC of 0.86, and at a consistency threshold of 0.90 the agent retained 49.4% of cases at 98.9% accuracy; blinded physician review of 181 cases agreed with the automated grading 92.3% of the time. The authors listed the limitations: retrospective simulation only, text only, one institution's data, lower accuracy in older patients and roughly five times the compute. Under that design, the remaining cases, those on which the agent's repeated runs disagreed, are routed to a person.
  • Evidence Jun 17, 2026
    Google's AMIE trial in Nature was a randomized, blinded virtual OSCE with 100 multi-visit cases and 21 primary care physicians; the model's management plans were rated appropriate 95% to 98% of the time versus 72% to 81% for the physicians. The same paper describes the system as "not ready for real-world translation," citing a text-only format with patient actors and no pharmacist or order entry. In April, Science published a Harvard study in which an OpenAI model matched attending physicians on emergency department triage and admission decisions using real records. Stanford researchers have repeatedly found that a physician using the model does no better than the model alone, and that the gap closes only when the physician commits to an assessment first and then compares.
  • Evidence Sept 11, 2026
    Becker's Hospital Review asked chief information officers and chief medical information officers at Cornell, Baptist Health, HSS, Premier Health, Christ Hospital, North Country Healthcare and Denver Health what the ADVOCATE agents of the Advanced Research Projects Agency for Health (ARPA-H) would need. Dr. Curtis Cole of Cornell said the agents will not succeed if the companies that build and oversee them have no liability. Joy Oh of Christ Hospital asked who is responsible when a supervisory agent misses an inappropriate recommendation, and Dr. Daniel Kortsch of Denver Health said the technology has been tested on curated cases, not safety-net populations. The consensus was narrow use cases, named accountability and escalation paths, not broad autonomy.