Is AI going to replace you? A map of where things stand, September 2026
This is the piece I wish someone had handed me a year ago. It is long because the question deserves it. The daily posts record what changes; this is the part that mostly stays put.
- The short version
- What the models can do, and what nobody has shown
- What runs without a doctor today
- The three ways physician work actually gets replaced
- Exposure by specialty
- The rules: FDA, Medicare, the states, and who gets sued
- Follow the money
- Recommendations
The short version
After several months of reading everything I could find, this is my summary. The models are now better than we are at the parts of medicine that fit in a text box. The best studies of 2026 say so, and I go through them below. Nobody has shown that a model running care on its own makes patients better off, mostly because almost nobody is allowed to try. And where physicians do get replaced, payment rules and liability law will decide it, not another benchmark.
Three points, in the order they worry me. The reasoning gap is real and growing. The only care in the United States that runs without a physician's signature is a short list: a retinal screening algorithm, one state's refill sandbox, and a cleared insulin-titration app that works inside a doctor's plan. And what actually reduces demand for physicians is not an autonomous AI; it is a medical assistant with a model and a remote doctor's signature, seeing four or five times the patients. That arrangement is legal everywhere and already running at scale.
Meanwhile the labor market says the opposite of the headlines. Physician pay rose again this year. Eighty-five percent of us say shortages affect our practice. Radiology, the specialty everyone said would be automated first, pays about $610,000. No health system in the country has cut physician positions and named AI as the reason; the cuts that have been reported are in IT and coding. What has happened instead is less dramatic and, I think, matters more: the time that ambient scribes save is being converted into higher productivity targets, and there is now decent evidence that leaning on the tools erodes skills.
What the models can do, and what nobody has shown
Start with the study that changed my mind. In June, Nature published Google's AMIE disease-management trial: a randomized, blinded, virtual OSCE with 100 three-visit cases across five specialties and 21 primary care physicians. AMIE's management plans were rated appropriate 95 to 98 percent of the time. The physicians' plans came in at 72 to 81 percent. Guideline alignment, treatment precision, drug reasoning: the model won every column. The authors then said, in their own paper, that the system is "not ready for real-world translation." Patient actors, text only, UK guidelines for non-UK physicians, no pharmacist, no order entry. Both statements are accurate, and the second one deserves as much weight as the first.
In April, Science published a Beth Israel and Harvard study in which an OpenAI model matched or beat attending physicians at triage, first-contact, and admission decisions on real Boston emergency department records. Adam Rodman, who ran it, was careful to say the results support teaming, not "AI doctor companies." OpenAI's own benchmark, HealthBench Professional, has GPT-5.4 at 59 against specialty-matched physicians at 43.7, with the physicians allowed the internet and unlimited time. OpenAI wrote the rubric and graded its own model, and its paper says the scores are "not directly comparable to real-world performance." Microsoft's MAI-DxO, the "four times more accurate than doctors" result from 2025, is still an unrevised preprint fifteen months later, with the physicians in the comparison denied both search tools and colleagues.
The most useful evidence is smaller and less exciting. The one large real-world deployment, OpenAI's work with Penda Health in Kenya across about 40,000 visits, cut diagnostic errors 16 percent and treatment errors 13 percent when the model acted as a safety net for clinicians who kept full authority. Patient-reported recovery did not change. And Stanford's series of trials on physicians using GPT-4 found something that should bother all of us: a physician with the model did no better than a physician without it, while the model alone scored 92 percent. Only when the workflow was redesigned so the physician commits to an assessment first and then compares did the combination beat physicians alone and roughly match the model. That is a finding about workflow more than intelligence: left to ourselves, we use these tools badly.
To sum up: on rubric-scored text cases, the model alone is at least as good as a physician alone. A physician with the model is often no better than the model alone unless someone thought about the workflow. And nobody has shown better patient outcomes from AI-led care, because nobody has been allowed to test it. On the physical, long-term, accountable parts of medicine, the question is untested rather than settled.
What runs without a doctor today
Shorter list than the press releases suggest. LumineticsCore diagnoses diabetic retinopathy with no clinician read and bills its own CPT code. Doctronic operates in a Utah regulatory sandbox that lets an AI renew about 190 non-controlled chronic medications, and after seven months it is still in Phase 1, which means a licensed practitioner signs every renewal. The AI recommended renewal 72 percent of the time and physicians agreed 91 percent of the time; Utah's own medical licensing board asked for the pilot to be suspended and was turned down. Legion Health got a parallel approval for fifteen psychiatric medications. UpDoc holds the first FDA clearance for a patient-facing large language model; it titrates diabetes medications by voice or chat, inside a plan a physician wrote. In Europe, Oxipit's ChestLink clears normal chest films on its own under a CE mark; nothing like it is cleared here.
Everything else that calls itself an AI doctor is a person plus a model plus a doctor's signature. Akido Labs says it is the first AI-native health system: one million patients, 100 clinics, medical assistants running AI-guided visits with a physician approving remotely, and a claimed four to five times the patients per physician. No outcomes trial has been published. Counsel Health offers a free AI front door and a $29 physician follow-up. K Health sits inside Mass General Brigham, Mayo, Penn, and Novant. Amazon gives Prime members up to five free messaging visits a year for more than 30 conditions. Hippocratic AI's voice agents have logged 250 million patient interactions doing intake, outreach, and post-discharge calls. Hims & Hers says AI now answers 80 percent of the questions from weight-loss patients on its new care platform. This is the part changing fastest, and it has more to do with labor costs than with robots.
The three ways physician work actually gets replaced
I find it clarifying to separate three pathways, because they move at different speeds and the loudest one is the slowest.
Full autonomy, meaning no human signature, is narrow and heavily regulated: retinal screening, refills in one state, protocolized titration. It grows one FDA clearance and one state sandbox at a time.
Supervised autonomy is the agent acting on its own with a physician reviewing exceptions or samples. This is what ARPA-H is now funding: $62.7 million for partially autonomous heart-failure agents that assess symptoms, adjust guideline-directed therapy, and order labs under a physician-prescribed plan, with FDA packages due within two years and a 2,500-patient trial at Kaiser. Baptist Health's chief medical information officer described it as a shift from human-in-the-loop to exception-based oversight. Expect to hear that phrase a lot.
Labor arbitrage is AI plus a medical assistant, nurse, or pharmacist plus a remote physician signature. It requires no new law and no new clearance, and it is the pathway that lowers the number of physicians needed per thousand patients. Akido is the proof of concept. The useful question about your own job is not whether a model can do it but whether a cheaper licensed person with a model and your signature could.
Exposure by specialty
| Work | Exposure | Why | What the data say |
|---|---|---|---|
| Radiology, pathology | High, slow | Image interpretation is the most benchmarkable task in medicine; autonomous normal-film clearing exists in Europe. | Pay about $610,000; exam volume up 31 percent from 2018 to 2024 with reads per radiologist flat; no US autonomous clearance. Bob Wachter calls it the easiest field for AI to replace, on a ten-to-fifteen-year horizon. |
| Telehealth refills, triage, low-acuity primary care | High, now | Protocolizable, text-native, price-sensitive. | Doctronic $39, Counsel $29, Amazon free messaging visits, Hims 80 percent AI-answered questions, Akido four to five times throughput. |
| Documentation, coding, prior auth, chart summarization | High, now | Already automated at scale. The open question is who keeps the time. | Abridge now sells pre-bill coding review; Epic's Ergo Visit drafts notes and orders; Stanford's discharge-summary agent had omissions in 25 percent of drafts. |
| Hospital medicine | Moderate | Cognitive and documentation-heavy, but bedside, high acuity, and accountable. Virtual night coverage is the substitution channel. | 64 percent of hospitalists use AI and a third of those users feel adequately trained (SHM, July 2026). Virtual hospitalist programs do everything but hands-on procedures. No AI-attributed ratio changes anywhere. |
| Family medicine, relationship care | Moderate, mixed | Inbox and refill work is exposed; the relationship is what patients say they are paying for. | 93 percent of non-mental-health Zocdoc bookings are in person, and Gen Z books in person 92 percent of the time; 90 percent of patients want a human escalation path; nobody has validated a bigger panel size. |
| Emergency medicine | Moderate-low | Triage models perform, but undifferentiated patients, procedures, and time pressure keep the human central. | Only 12 percent of emergency physicians worry AI will reduce EM headcount (ACEP, February 2026); ED scribes did not shorten waits. |
| Procedures, surgery, bedside critical care | Low | Physical, real-time, accountable. | Microsoft's occupational AI-applicability score for healthcare practitioners is 0.13; hands-on health roles score 0.04 to 0.06, among the lowest in the economy. |
| Psychiatry and therapy | Moderate, contested | Chatbots are capable and cheap; states are restricting them. | Twelve chatbot laws in eleven states this year; Illinois and Nevada ban AI therapy outright; Pennsylvania sued Character.AI for practicing medicine. |
Hospital medicine is my day job, so a word about it. Within eighteen months at any Epic hospital, expect drafted notes and drafted orders waiting for your signature, agent-written discharge summaries, automated first-pass utilization review with the physician advisor reserved for the hard cases, and pressure toward virtual nights. None of that reduces the number of hospitalists a hospital needs on days, but it will raise the census you are expected to carry, and it moves liability onto whoever signs. So read what you sign.
The rules: FDA, Medicare, the states, and who gets sued
FDA
In January the agency loosened the two guidances that matter most to a practicing physician. Under the revised clinical decision support guidance, software that hands a clinician a single clinically appropriate recommendation, including risk scores and differential-diagnosis tools, is not a regulated device; time-critical uses like stroke and sepsis, image analysis, and anything that acts without a clinician still are. The revised wellness guidance lets wearables report blood pressure, glucose, and lactate as wellness data so long as they make no disease claims. The AI-enabled device list passed 1,500 authorizations. Then on August 18 the FDA published a discussion paper on generative-AI devices proposing "competency-based" evaluation: test the model the way you would test a clinician, then monitor it after market. Comments close October 19. Its digital health center says formal guidance is coming. And in July, ten companies, including Anthropic, Microsoft, Amazon, Doctronic, and Hippocratic, demonstrated "AI doctor" technology to FDA, CMS, and HHS officials in a closed-door session at White Oak.
Medicare
Two programs matter. ACCESS, which launched in July, pays outcome-linked amounts for technology-enabled care of hypertension, diabetes, musculoskeletal pain, depression, and anxiety, with heart failure, COPD, substance use, and tobacco added next spring. About 160 organizations are in, including Headspace, Noom, and WeightWatchers. Paired with the FDA's TEMPO pilot, unauthorized AI tools are already treating Medicare beneficiaries inside it under enforcement discretion. The other program is WISeR, AI-assisted prior authorization in six states, and it is the cautionary tale: records obtained by the Electronic Frontier Foundation show a vendor that told CMS it would auto-affirm requests because it was not ready, two vendors that denied more than 20,000 requests in three months, vendors paid per denial, and one request that sat for 83 days. A Senate attempt to kill it failed 46 to 50 in July. If you are fighting a prior authorization anywhere, this is your evidence that AI gatekeeping fails in practice.
The proposed 2027 fee schedule cuts the conversion factor 1.19 percent, turns G2211 into a percentage modifier, restricts remote monitoring to practice-employed staff (which would end vendor-run RPM in Medicare on January 1), and asks for input on ambient documentation and AI-enabled wellness visits. The final rule is expected around November 1. Telehealth flexibilities run through December 31, 2027. The DEA's telemedicine prescribing rule runs through the end of this year. And on September 14 the New York Times reported that officials have discussed a Medicare payment category for company-run "AI physicians" at 60 to 80 percent of the human rate. Nothing has been proposed yet, but of everything I follow on payment, this is the item I check first.
The states
Twenty-two states enacted 33 health-AI laws in the first seven months of 2026, on top of last year's wave (Manatt keeps the best tracker). The pattern is remarkably consistent: tell patients when AI is used, keep a licensed human accountable for clinical use, require human review of payer denials, and stop chatbots from posing as mental-health providers. Colorado, Texas, Maine, Vermont, and Louisiana now put the duty to review AI output on the practitioner in statute. Alabama, Georgia, and Minnesota require human review of payer utilization decisions. Eleven states passed chatbot laws. California has a bill on the governor's desk, AB 1979, that would bar AI from performing licensed clinical work or directing unlicensed staff to do it, which is aimed squarely at the labor-arbitrage model; he has until September 30. No state has authorized an AI to practice medicine, and the Federation of State Medical Boards said in August that AI "is not ready to be independently licensed like a physician." About twenty liability bills were introduced. Zero passed.
Who gets sued
Liability still lands on the licensed human, and in at least five states that is now written down. No malpractice suit over an AI scribe has been filed yet, but the malpractice carriers are warning about the failure modes: hallucinated template text, omitted non-verbal findings, and the slow decay of how carefully you read a draft once you trust it. The plaintiffs' bar is testing a different theory against the model companies: unlicensed practice of medicine. Winters v. OpenAI, filed in July, alleges ChatGPT told a man with unstable blood pressure to stay in his recliner and he developed a pulmonary embolism. Pennsylvania sued Character.AI under its Medical Practice Act after a chatbot posed as a licensed psychiatrist. The AMA's June policy says AI "cannot replace physician judgment," and its CEO wants the companies to "take on higher levels of liability." My practical read: a practice that documents physician review of every AI output and tells patients when AI is used is compliant everywhere today. One that lets AI output reach a patient unreviewed is exposed in at least five states and, I suspect, in front of any jury.
Follow the money
AI took 46 percent of healthcare venture dollars in 2025. Rock Health counted $7.4 billion in US digital health funding in the first half of 2026, with twenty megadeals taking 45 percent of it and zero IPOs. The highest valuations sit on two kinds of company. The first strips administrative labor from existing physicians: OpenEvidence at $12 billion (free to clinicians, ad-supported, and used daily by more than 40 percent of US physicians by its own count; it considered, then shelved, a round at $20 billion in July), Commure at $7 billion, Abridge at $5.3 billion. The second sells consumers a longevity subscription that needs a physician to interpret it: Function Health at $2.5 billion took another $450 million in July, and Oura filed for a Nasdaq IPO in September with $1.4 billion in trailing revenue. The annual lab bundle is converging on $300 to $600: Function's starts at $365 and Fountain Life's new tier is $595. Labs and scans are becoming commodities; the margin is in interpreting them and answering for the result.
In public markets the winners are toll collectors on physician workflow: Doximity's shares jumped by roughly two-thirds after a quarter in which its AI scribe users grew tenfold; HeartFlow keeps raising guidance on AI plaque analysis; Tempus won ARPA-H money for autonomous cardiology. Undifferentiated telehealth is losing: Teladoc's revenue fell and BetterHelp shrank. Revenue-cycle software is now priced as an AI casualty even when it uses AI itself; Waystar hired bankers to explore a sale after its stock fell 24 percent this year. The one ETF that calls itself a pure play on healthcare AI holds genomics and surgical robots, not care delivery. None of this is investment advice; it is a map of where investors think physician time is worth paying for and where they think it can be replaced.
The bear case for physicians is worth reading in full. Ezekiel Emanuel and his co-authors, Vinod Khosla among them, argued in JAMA in August and again in STAT on September 9 that by 2030 autonomous AI will beat AI-assisted physicians at history-taking, differential diagnosis, test selection, guideline-concordant prescribing, and chronic disease management, and that value will move to procedures, presence, accountability, and relationships. The AMA's CEO answered in STAT the same day. What strikes me is that both sides land in the same place: the physical, relational, and accountable parts of the job are the parts that stay.
Recommendations
I am not going to tell you to "learn AI." Everyone says that and it means nothing. What follows is specific to where you work.
If you are a hospitalist. Sign nothing you have not read; in five states that duty is now law, and everywhere it is your license. Ask your group, in writing, how productivity targets will change as ambient and agent tools roll out, because Doximity's data say the targets rise and the pay does not, and the group that has that conversation early sets the terms. Get on the AI governance committee or the Epic agent rollout. Two-thirds of hospitalists who use these tools say they were not adequately trained on them; the person in the room when these tools are configured is the person they end up serving.
If you are in primary care. The refill, inbox, and triage work is going to a model plus a nurse whether you like it or not, and you should let it, because it was never the part patients valued. What they are paying for, and what the surveys say they want, is a human who knows them, verifies what the machine says, and is accountable. Nobody has validated a bigger panel size with AI; do not let anyone hand you one without evidence. If you own your practice, you keep the AI dividend. If you are employed, you are negotiating for it.
If you are in the emergency department or do procedures. Every exposure index in existence bottoms out on hands-on, time-critical, physically present work. The virtual hospitalist companies stop at procedures. One in thirteen American emergency departments has no attending physician on site around the clock, and nearly nine in ten of those are critical access hospitals. The rural hospitals about to receive $50 billion of new federal money are being told AI will save them, and their leaders don't believe it. If you have been thinking about airway, access, or rural coverage, the market is telling you to do it.
If you are thinking about the business side. Every company that wants to run AI-assisted care legally needs a physician in the loop, and they pay for it: part-time medical director roles run $3,000 to $10,000 a month, model-evaluation work pays $130 to $180 an hour for internists and emergency physicians, advisory boards pay cash and equity, and the first AI-related malpractice cases are being filed now, which means expert witnesses with real AI-governance experience will be in demand by 2028. Chief medical AI officer roles pay $420,000 to $720,000 base and require an MD in nine out of ten searches. Recruiters want deployment stories, not certificates. The clinical informatics board's practice pathway closed last year; if you want a credential, Stanford's Coursera specialization is cheap and Harvard Chan runs a $2,600 online course in December, but building and governing one real tool in your own practice will teach you more and prove more.
How to use the tools without losing your edge. The evidence on this is uncomfortable and I take it seriously. Experienced endoscopists' adenoma detection fell from 28 to 22 percent after three months of AI assistance. Pathologists reversed correct diagnoses more than 30 percent of the time when a wrong AI prompt was on the screen. So: commit to your own assessment before you look at the model's, then compare; the Stanford data say that is the workflow that closes the gap with the model. Ask what would harm the patient if it were left out, because these systems fail by omission far more than by commission. Keep deliberate AI-off intervals for reads and procedures and track your unaided numbers. Read every draft as if you wrote it, because legally you did. And never hand an irreversible, high-stakes decision to anything that cannot be deposed.
The Friday letter: the week that mattered and specific recommendations, free, in your inbox.
Get the Friday letter