When the model is free, the proof is the product

Open models now outscore most radiologists on paper, and the FDA has refused to let a vendor's track record stand in for review. Radiologists and pathologists should claim the job of proving, on their own patients, that a model works, and be paid for it.

In 2016 Geoffrey Hinton, the computer scientist, said radiologists were like a coyote already over the edge of the cliff and that we should stop training them. Ten years on, the past ten days brought two results that would have pleased him. Alibaba released a model that reads contrast abdominal CT for 146 findings and outperformed 23 of 26 radiologists, with its code free to anyone. Nature Medicine published EAGLE, which found esophageal cancer on plain CT better than all 17 radiologists in its reader study. The same fortnight, SullivanCotter reported that radiology pay is up 12 to 18 percent since 2024. The coyote is still in the air, and it has had a raise.

The raise is not a mystery. Reading was never the whole job, and the part that remains is becoming the scarce one: knowing whether a particular model is right about the particular patients in front of you. When a capable reader costs nothing, the value moves to the proof. Radiologists and pathologists should claim that work formally, as the owners of local validation and monitoring for every model that touches their images, with the credentials, protected time and pay that go with it. The alternative is to leave it to vendors, to hospital IT committees, or to an FDA clearance granted once, on other people's patients.

Look first at how far this week's headline results sit from an American reading room. EAGLE's numbers are superb on paper: 90 percent sensitivity and 98.5 percent specificity on an external test set of 11,466 patients. Run prospectively, its flags had a positive predictive value of 42 percent, so most of them were wrong. Nearly all its patients were Chinese, in whom squamous cell cancer dominates; in the United States the common tumor is adenocarcinoma of the distal esophagus, which the model has not been shown to find. It performed worse in women because the training set skewed male. Alibaba's model was tested on Chinese patients only, on contrast CT only, with no prospective trial. None of this faults the science; it describes the ordinary distance between a validation set and a hospital.

That distance is measurable here, too. Northwell Health ran an FDA-cleared aneurysm detector on 3,856 consecutive CT angiograms and published the result in JACR this month. The tool found 55 aneurysms the radiologists missed, a 39 percent relative gain, but its positive predictive value was 78 percent against the radiologists' 93. The ratio of useful finds to false alarms was 2.57 on inpatients, 1.00 in the emergency department and 0.67 in outpatients, where false alerts outnumbered true ones. One cleared product gave three different answers depending on which door the patient came through. The American College of Radiology's own registry paper, covering 62,835 studies at 16 facilities, notes that real-world performance often falls short of the figures shown for authorization. And Epic's end-of-life index, tested last week at 39 hospitals, ranked patients well while overstating how many would die.

The FDA understands this, and on September 17 it said so. Its final order refused Harrison.ai's request to let experienced makers of radiology detection and triage software skip 510(k) review. In its earlier denial letter the agency argued that a radiologist cannot be counted on to catch a faulty algorithm, because the truth about an image may not be available before the error does its harm. That is correct, and it is also a confession of the limits of premarket review, which happens once, on someone else's data, before the model meets a single local scanner or referral pattern. Clearance tells you a device was safe and effective somewhere. Whether it helps your inpatients and harms your outpatients is a question only local data can answer.

At the moment that question is being answered, when it is answered at all, by information officers. UW Health sorted more than 100 AI features that arrived in a single software upgrade into risk tiers, and its governance review, rather than the departments using the tools, decides which are switched on. Rush's chief information officer, Jeff Gautney, moved a five-person chart-abstraction job to an AI vendor and said the job is not coming back. These are sensible people doing a job nobody else claimed. Radiology already has the model for claiming it. No department lets a new CT scanner image a patient on the strength of its clearance; a medical physicist runs acceptance tests first and repeats them every year. An algorithm that changes what gets reported deserves at least the scrutiny given to the scanner's dose.

The strongest objection is cost. Local validation takes data, statisticians and time that a community group reading for a critical-access hospital does not have. Demanding it would favor academic centers, slow the tools that are already catching missed aneurysms, and widen a gap Harrison.ai puts at AI touching about 1 percent of American diagnostic radiology. The Northwell study used six weeks of consecutive exams, the scale of a departmental quality project. Shared infrastructure exists: the ACR's Assess-AI registry pools de-identified outputs and reports, and North Carolina has just put $4.4 million into a network to help rural hospitals choose and monitor AI. The cost of not checking is paid anyway, in outpatient false alarms and in adenocarcinomas a Chinese-trained model was never taught to see.

Pathology is at the same threshold. The CPT Editorial Panel's September agenda included a Category III code for algorithmic analysis of digitized prostate slides, and a rewrite of the pathology code application that would ask applicants about their software and its validation. That vocabulary, once written, will decide what is paid for. The logic reaches beyond the image specialties: the cardiologist asked to act on an algorithmic STEMI call, the dermatologist handed an impedance reading of a lesion, and the generalist reading a report that a model helped write all have a stake in whether someone local has checked the tool.

Radiology and pathology chairs should adopt a written rule this month: no model contributes to a report until the department has run it on a consecutive local sample and published sensitivity, positive predictive value and results by care setting, with a named physician responsible for each tool. Medical staff offices should treat switching on a clinical algorithm the way they treat granting privileges, with review, a term and renewal. Radiology groups negotiating hospital contracts should write validation and monitoring into the service agreement as paid work, and refuse terms that let a software upgrade switch on reading tools without the department's sign-off. The ACR and the College of American Pathologists should use the FDA's comment period, which closes October 19, to ask that postmarket monitoring rest on local performance data, and should press the CPT panel and CMS to recognize validation as physician work. A free model that beats 23 of 26 radiologists still needs a 27th to say whether it works here.

The MAHA Summit meets in Washington on Sept. 29 with OpenAI, Anthropic, Hims & Hers and Sword Health among its sponsors. Gov. Gavin Newsom has until Sept. 30 to act on four California health-AI bills: AB 1979 and AB 2575, on AI directing unlicensed staff and the right to override AI, and SB 903 and SB 503, on AI therapy and bias monitoring of decision support. The third ACCESS cohort starts Oct. 1. Brigham and Women's nurses voted 93% on Sept. 24 to authorize an open-ended strike, which requires 10 days' notice.

Comments on the FDA's generative-AI device discussion paper are due Oct. 19, and the HIMSS AI in Healthcare Forum meets in San Diego Oct. 22 to 23. The 2027 physician fee schedule final rule, including the remote monitoring provisions, is expected around Nov. 1, and RSNA opens in Chicago on Nov. 29.