Free advice

AI will soon give clinicians in much of the world a specialist's knowledge for free. Poor and rural patients will gain only if someone supplies the tests, drugs, specialists and money it calls for.

A rope footbridge across a gap between two rock ledges, half of its planks ordinary gray boards and the other half glowing amber light.

The most useful thing artificial intelligence will do for medicine in the next few years is to make the clinician who works alone a little less alone. Last month OpenEvidence, a tool that answers doctors' clinical questions from the medical literature, said it was rolling out a free version in about 100 low- and middle-income countries, from Uganda and Sudan to Haiti and Mongolia, with Anthropic providing its AI models, engineers and money. The company says that on an average day, most American doctors use it.

That is a real gain, and it will matter most where specialists are scarce: rural America has about a third as many doctors per person as its cities. But knowledge is the cheapest thing a sick person needs. What rural and poor places lack is what comes after the answer: the scan, the drug, the specialist who will take the referral, the road to reach them and the money to pay for it all. Many lack even electricity: close to a billion people in poorer countries depend on clinics with unreliable power or none. AI will widen access to care only where it is built into those things. Treating it as a substitute for them is the magical thinking that Silicon Valley keeps selling.

Some gains are already measurable. In a trial in Mayo Clinic's primary care practices, teams shown an AI reading of routine electrocardiograms diagnosed nearly a third more patients with a weak heart pump. Young people with diabetes offered an eye exam graded on the spot by AI all completed it; of those sent to an eye doctor, about a fifth did. In Zambia and North Carolina, novices using a cheap ultrasound device with built-in AI dated pregnancies as accurately as trained sonographers. The World Health Organization lets software stand in for human readers of chest X-rays when screening for tuberculosis. Duller technology helps too: when specialists answer primary care doctors' questions in writing, waits fall from months to days or weeks.

Money is following. The Gates Foundation and OpenAI have committed $50 million to bring AI to 1,000 primary care clinics by 2028, starting in Rwanda. OpenEvidence says it will tailor the tool to local guidelines, and Penn Medicine plans to design tools from the ground up with clinicians in Botswana. Patients are not waiting: OpenAI says Americans living more than half an hour's drive from a hospital send ChatGPT more than half a million health messages a week.

The first rigorous tests are sobering. A randomized trial in 16 Kenyan clinics found that an AI assistant built on OpenAI's GPT-4o was safe and helped clinicians reach appropriate diagnoses. It did not reduce the share of patients whose treatment failed. An earlier study at the same chain of clinics, paid for by OpenAI, had found 16% fewer diagnostic errors as judged by reviewing doctors, but no clear difference in whether patients felt better. In Britain, people given chatbots to work through medical scenarios did no better than people left to their own devices, though the chatbots alone usually named the right condition. Studies of OpenEvidence itself are few and small: a review of eleven found its answers generally relevant and backed by evidence, but also that it often reinforced decisions doctors had already made, and a disputed study rated its answers to doctors' real questions no better than Google's AI summaries.

The trouble is what comes after the answer. In the Mayo trial, about half the patients the AI flagged did not get an echocardiogram to confirm it. In the eye trial, more than a third of those with abnormal results had not seen an eye doctor six months later. When Google put its eye-screening AI into clinics in Thailand, a fifth of the images were rejected, slow internet delayed the uploads, and patients whose images failed were told to see a specialist elsewhere on another day. More than half of rural American counties have no hospital that delivers babies, and no chatbot can perform a cesarean section.

None of this dims the optimists. Vinod Khosla, a venture capitalist, predicted in 2012 that computers would replace 80% of what doctors do. Bill Gates said last year that within a decade great medical advice would be free and commonplace. Sam Altman, OpenAI's chief executive, has called ChatGPT a better diagnostician than most doctors in the world, though he said he did not want to entrust his own care to it without a human doctor. Mehmet Oz, who runs Medicare and Medicaid, has said that the best way to help some rural communities will be AI-based avatars; his agency later explained that he meant tools that extend clinicians' reach, not replace them. According to The New York Times, officials have discussed paying AI doctors run by technology companies as much as 60% to 80% of what human doctors earn.

The record should temper them. MD Anderson Cancer Center abandoned a Watson project after spending $62 million, and internal IBM documents cited unsafe and incorrect treatment recommendations from its cancer software. Babylon Health claimed in 2018 that its chatbot beat the pass mark on practice questions for Britain's exam for family doctors, and it ran Rwanda's national telemedicine service, with nearly four million consultations, until the company collapsed in 2023 and the service closed. Care that depends on a company's balance sheet ends when the balance sheet does. In April OpenEvidence itself stopped serving doctors in the European Union and Britain, citing regulatory uncertainty.

There are quieter risks too. Doctors defer to confident software: in a trial in Pakistan, physicians whose AI assistant was wrong in half the cases scored 14 points lower on diagnostic reasoning than those whose assistant was right, despite 20 hours of training in AI. Dermatology software tested on images of diverse skin tones did worse on dark skin. Chatbots can falsely reassure: given written scenarios describing emergencies, ChatGPT's health service under-triaged half, advising some to be seen within a day or two. OpenEvidence is supported by advertising, much of it from drug companies. And cheap software may become the only care that poor and rural patients are offered. "I'd be curious if Dr. Oz would want an avatar treating his own family," said Carrie Henning-Smith, who studies rural health at the University of Minnesota.

The strongest objection is that something beats nothing. Where there is no doctor for a hundred miles, an imperfect assistant is better than none, and insisting on trials first delays help for people who need it now. Where the alternative really is nothing, that is often right, and the Kenyan trial found its assistant safe. But the evidence so far shows that these tools help when they are tied to care, and expert care for places without experts is just what Watson and Babylon promised. Trials need not be slow: the Kenyan one enrolled nearly 10,000 patients in under three months. Patients in rural Uganda and rural Montana deserve the same proof as patients in Boston.

What would let AI widen access is mostly not AI. Governments should finish wiring and powering rural clinics. Medicare, Medicaid and health ministries should pay for, and keep paying for, the links that turn advice into care: specialist advice by message, video visits and the follow-up the software recommends. They should judge AI tools by completed referrals and patients' health rather than by use, and pay for software that acts on its own only after trials show it helps. Companies giving tools away should publish what happens to patients, build for local drugs, tests and weak connections, say how long free will last and keep advertising out of clinical tools. Big health systems should share the work of choosing and checking tools with small hospitals, as a new network in North Carolina means to. And clinicians should argue with the machine, making their own assessment and comparing it with the software's, an approach that helped doctors in trials led by Stanford researchers where simply handing them the tool did not.

A family doctor in Montana and a clinician in Uganda will soon carry a specialist's knowledge in their pockets. The advice is becoming free. The care it calls for still has to be paid for, staffed and built.

  • Reuters reported on Sept. 22, 2026, that OpenEvidence and Anthropic are rolling out a specialized version of OpenEvidence's clinical decision-support tool, free to health care providers, in about 100 countries, including Uganda, Angola, Sudan, Haiti and Mongolia, according to a list provided by OpenEvidence; financial terms were not disclosed. Fierce Healthcare reported on Sept. 28 that Anthropic is providing Claude credits, dedicated engineers and funding for a team focused on Africa and Asia, and that OpenEvidence will tailor the platform to local medical guidelines. Anthropic's president, Daniela Amodei, told Reuters that "the market incentives alone would not let it happen without this type of entrepreneurial, philanthropically minded work." The reports did not say how clinicians will be verified, which languages will be supported or whether the free version will carry advertising. Penn Medicine announced on Sept. 15 that it will offer OpenEvidence to more than 10,000 clinicians and build customized tools with clinicians of the Botswana-UPenn Partnership; OpenEvidence began a pilot with 45 Rwandan clinicians in 2025.
  • OpenEvidence says it is used daily, on average, by the majority of U.S. physicians, that U.S. clinicians consulted it 42 million times in August 2026 and that 1.12 million verified U.S. clinicians, including nurses, nurse practitioners and physician assistants, use it (company figures). It is free to verified U.S. clinicians and supported by advertising, many of the ads from drug companies, shown on roughly 5% of queries (Becker's, June 2026). It raised $250 million at a $15 billion valuation in September 2026 (Forbes). In late April 2026 it stopped serving users in the European Union and the United Kingdom, citing regulatory uncertainty including the EU AI Act. A June 2026 Nature Medicine study by NYU Langone researchers scored 100 real clinical queries on a 4-point scale: Gemini 3.62, ChatGPT 3.54, Claude 3.52, Google's AI Overview 3.27, OpenEvidence 3.24 and UpToDate's AI 3.17; the company asked the journal to retract the paper, and a published critique said its real-world evaluation overstated precision. An August 2026 systematic review of 11 studies found that it "generally generated clinically relevant, evidence-supported responses" with fewer fabricated citations than general-purpose models, that it "often reinforced rather than altered clinical decisions," and that prospective real-world studies are needed.
  • In a study funded by OpenAI and posted as a preprint in July 2025, clinicians at Penda Health clinics in Nairobi who had an AI assistant (AI Consult) made 16% fewer diagnostic errors and 13% fewer treatment errors, as judged by physician reviewers, across 39,849 visits at 15 clinics; the share of patients who said they were not feeling better did not differ significantly (3.8% vs. 4.3%), with about 60% of outcome data missing. A pragmatic cluster-randomized trial of a GPT-4o version in 16 Penda clinics (Nature Medicine, June 26, 2026; 9,691 patients enrolled April 22 to July 16, 2025; 103 clinical officers) found treatment failure within 14 days in 2.2% of patients with the assistant and 2.0% without (adjusted odds ratio 0.77, 95% confidence interval 0.55 to 1.08); appropriate diagnosis improved (adjusted odds ratio 1.74). The authors concluded that the assistance "was safe but did not reduce treatment failure within 14 days and any benefit, if present, is probably modest." The trial was funded by the Gates Foundation.
  • EAGLE (Nature Medicine, 2021): 120 primary care teams from 45 Mayo Clinic clinics or hospitals and 22,641 adults; access to AI readings of routine ECGs raised new diagnoses of low ejection fraction from 1.6% to 2.1% (odds ratio 1.32), and 49.6% of AI-positive patients in the intervention arm had an echocardiogram. Nigeria (Nature Medicine, 2024): an AI-enabled digital stethoscope doubled detection of left ventricular systolic dysfunction in pregnant and postpartum women (4.1% vs. 2.0%). ACCESS (Nature Communications, 2024): among young people with diabetes in Baltimore, all of those offered a point-of-care autonomous AI eye exam completed it, against 22% of those referred to eye care; 64% of those with abnormal AI results saw an eye care provider within six months. JAMA (2024): novices in Lusaka, Zambia, and Chapel Hill, N.C., using a low-cost AI-enabled ultrasound device estimated gestational age as accurately as credentialed sonographers (mean absolute error 3.2 vs. 3.0 days). WHO (2021): computer-aided detection software "may be used in place of human readers" of digital chest X-rays for TB screening and triage in people 15 and older (conditional recommendation, low certainty of evidence).
  • OpenAI says more than 230 million people ask ChatGPT health and wellness questions every week, and that ChatGPT averaged more than 580,000 health care messages a week from U.S. "hospital deserts," places more than a 30-minute drive from a general hospital (company figures, January 2026); its Health feature opened to U.S. adults on July 23, 2026. In a randomized study of 1,298 members of the public in the U.K. (Nature Medicine, Feb. 9, 2026), language models tested alone identified the relevant conditions in 94.9% of cases, but participants using them identified relevant conditions in fewer than 34.5% of cases and the right course of action in fewer than 44.2%, no better than a control group. Given 60 clinician-written cases (Nature Medicine, Feb. 23, 2026), ChatGPT Health under-triaged 52% of gold-standard emergencies, directing some to evaluation in 24 to 48 hours rather than the emergency department.
  • In 2022 there were 286 active physicians per 100,000 people in urban U.S. areas and 98 in rural areas, with nearly 7 times as many cardiologists and more than 8 times as many dermatologists and gastroenterologists in urban areas (AAMC). In 2022, 59% of rural counties lacked hospital-based obstetric services (HHS, 2024). The Sheps Center counts 197 rural hospital closures and conversions since January 2005, not counting 56 hospitals now operating as Rural Emergency Hospitals. The American Hospital Association says 48% of rural hospitals operated at a loss in 2023. Worldwide, WHO projects a shortfall of 11 million health workers by 2030; about 4.6 billion people lacked coverage for essential health services as of 2023; close to 1 billion people in low- and lower-middle-income countries are served by health facilities with unreliable electricity or none; and 2.2 billion people remain offline (ITU, 2025).
  • In Thailand, Google's diabetic retinopathy system, observed in 11 clinics, rejected 393 of 1,838 images (21%) as ungradable in its first six months, often because of poor lighting; slow connections delayed uploads; and patients with rejected images were told to see a specialist at another clinic on another day, until Google changed the protocol so that specialists reviewed ungradable images (CHI 2020; Google). In a large U.S. county health system, 28% of patients referred after diabetic eye screening kept a first ophthalmology appointment, were recommended for treatment and started it (BMJ Open Diabetes Research & Care, 2020). A 2026 meta-analysis of six studies found that AI-assisted screening raised referral uptake (relative risk 1.89), most where the referral pathway was redesigned.
  • MD Anderson Cancer Center canceled its Watson project in 2016 after spending $62 million (IEEE Spectrum). STAT reported in July 2018 that internal IBM documents cited "multiple examples of unsafe and incorrect treatment recommendations"; IBM said it had learned and improved Watson Health from client feedback. IBM agreed in January 2022 to sell its Watson Health data and analytics assets to Francisco Partners. Babylon Health said in June 2018 that its AI scored 81% on test questions modeled on the MRCGP exam, against a 72% average pass mark, a claim the Royal College of General Practitioners disputed. Its Rwandan service, Babyl, ran the country's first nationwide telemedicine service from 2019 to September 2023, with 3.90 million consultations; Babylon's U.S. subsidiaries filed for Chapter 7 bankruptcy in August 2023, and the Rwandan service was wound down.
  • In a randomized study of 44 physicians in Pakistan with 20 hours of AI-literacy training (NEJM AI, April 2026), those whose GPT-4o recommendations were wrong in half of six cases scored 73.3% on diagnostic reasoning against 84.9% for those given error-free recommendations (adjusted difference -14.0 percentage points). In four Polish centres, experienced endoscopists' adenoma detection rate without AI fell from 28.4% to 22.4% after AI was introduced (Lancet Gastroenterology & Hepatology, 2025; observational). Dermatology AI models performed substantially worse on dark skin tones in the Diverse Dermatology Images set (Science Advances, 2022). A preprint found that advertising shifted general language models' choice of an advertised drug by 12.7 percentage points on average when two drugs were both appropriate; OpenEvidence was not tested. Carrie Henning-Smith of the University of Minnesota Rural Health Research Center, on the CMS administrator's comments about AI avatars: "I don't like the idea of rural populations being treated as guinea pigs."
  • Vinod Khosla wrote in Fortune in December 2012: "Eventually, computers will replace 80% of what doctors do and amplify their capabilities." Sam Altman told a Federal Reserve conference on July 22, 2025 that ChatGPT is "like a better diagnostician than most doctors in the world," and added: "I really do not want to like entrust my medical fate to ChatGPT with no human doctor in the loop." Mehmet Oz said at an event in Washington on Feb. 2, 2026 that "the best way to help some of these communities" would be "AI-based avatars" (NPR); CMS said he meant tools that extend clinicians' reach. The New York Times reported on Sept. 14, 2026, citing one person involved, that officials have discussed paying AI doctors 60% to 80% of what human doctors earn. On the other side of the ledger: Medicare has paid for interprofessional (eConsult) codes 99451 and 99452 since 2019; Congress extended Medicare's telehealth flexibilities through Dec. 31, 2027; the North Carolina Collaborative Health AI Network, launched Sept. 22, 2026 with a $4.4 million Duke Endowment grant, will help rural and resource-limited hospitals choose and implement AI; and the Gates Foundation and OpenAI committed $50 million in January 2026 to reach 1,000 primary care clinics by 2028, starting in Rwanda.