Illustration by tuput
English
At 16 Penda Health clinics, clinicians with a GPT-4o assistant built into their records system had treatment failures in 2.2% of patients, against 2.0% without it. The independent panel still scored the assisted notes, diagnoses and treatment plans as better.
A randomised trial in Kenya found no statistically significant drop in treatment failures when clinicians used an AI assistant built into their patient records. Within 14 days of a visit, 2.2% of patients seen by assisted clinicians had a treatment failure, against 2.0% of patients seen by unassisted ones. The paper appeared in Nature Medicine on 26 June 2026, and its authors conclude that the tool was safe and that any benefit is probably modest.
The same trial found the assistant improved something else: the quality of what clinicians wrote down and decided, as judged by a panel of physicians.
What was tested and how
The tool is called AI Consult. It runs OpenAI’s GPT-4o (the May 2025 release) inside the electronic medical record used by Penda Health, a network of primary care clinics in Nairobi and Kiambu counties. While a clinician writes up a visit, it reads the notes and flags problems with a green, yellow or red alert. Clinicians did not have to follow it, and patients never saw its screen. The arrangement is closer to a surgical robot that a human drives than to an autonomous doctor: the clinician stays responsible for every decision.
It was a pragmatic cluster-randomised trial: it ran inside routine care, and the units randomised were clinicians, not patients. Sixteen Penda facilities took part. A total of 103 clinical officers, the clinicians who saw the patients, were split into 52 using the tool and 51 using the same record system with the tool switched off. Enrolment ran from 22 April to 16 July 2025.
The trial is registered in the Pan African Clinical Trials Registry (PACTR202502499779176), and its protocol, version 2 dated June 2025, is public on Zenodo, an open research archive. PATH, the global health nonprofit that sponsored the study, announced the trial on 11 March 2025 and expected results by the end of 2025.
The primary result
The prespecified main outcome was treatment failure within 14 days of enrolment. It was a composite judged by a panel of six family physicians, with two reviewing each case and a third settling disagreements. A failure counted if the patient came back with unresolved symptoms, had an unplanned escalation to higher-level or emergency care, or had a safety event such as a missed diagnosis, an inappropriate prescription or a death. Follow-up was by phone, on days 3 and 14.
The abstract says 9,691 patients were enrolled, out of 17,626 screened. After exclusions, including 254 encounters with protocol non-adherence and 90 patients lost to follow-up, 9,347 were analysed. The counts were 102 failures among 4,693 patients in the assisted arm (2.2%) and 94 among 4,654 in the control arm (2.0%). The adjusted odds ratio was 0.77, with a 95% confidence interval of 0.55 to 1.08 and a P value of 0.13.
The raw counts and the adjusted figure point in opposite directions. Raw, the assisted arm had slightly more failures, while an odds ratio below 1 favours the assisted arm. The adjusted figure comes from a model that allows for clustering by clinical officer and facility, and the paper does not reconcile the two. Adding patient sex, age, time of day and day of week gave an odds ratio of 0.72 (95% CI 0.50 to 1.03, P = 0.07). NPR reported the result as a 23% decrease in treatment failures that was not statistically significant. An odds ratio is a different measure from a percentage drop in failures.
The authors translate the interval into plain terms in their Discussion: somewhere between 13 fewer and 1 more treatment failure per 1,000 patients. Their model-based estimate of the risk difference was 0.5 percentage points lower in the assisted arm, with a 95% credible interval from 1.3 points lower to 0.1 points higher.
Where the AI did better
A reviewing panel checked 2,000 encounters for documentation quality. Assisted clinicians were more likely to record an appropriate diagnosis (adjusted odds ratio 1.74, 95% CI 1.28 to 2.36), write a full note (1.68, 1.24 to 2.27) and set out an appropriate treatment plan (1.71, 1.25 to 2.34). All three had P values below 0.001.
Other measures were flat. Among 826 patients surveyed, satisfaction did not differ between arms, and median consultation time was 11 minutes in both. A post hoc look at spending, which the authors flag as not confirmatory, found the tool cost about $0.04 per patient to run and that mean antibiotic spending was $3.71 in the assisted arm against $3.85 in the control arm.
Experts also reviewed 1,000 of the assistant’s red alerts. They rated 49.4% definitely safe and 42.4% mostly safe. About 3.1% were somewhat unsafe or inappropriate and 1.1% unsafe and inappropriate. Clinicians followed red alerts fully in 19.5% of cases, partly in 57.3% and not at all in 23.2%.
Why a null result may not mean no effect
The trial was sized to detect a halving of treatment failures, from 2% to 1%, using about 9,000 encounters. The authors say it was powered for a larger effect than the one observed. In the Discussion, they write that post hoc simulations suggested detecting modest differences in rare outcomes would take more than 100,000 patients. Dr Bilal Mateen, a co-author and PATH’s chief AI officer, told NPR the figure is about 139,000.
An earlier Penda study, posted to arXiv in July 2025 as a quality-improvement study, looked at 39,849 visits across 15 clinics. Independent physicians rated visits for clinical errors and found 16% fewer diagnostic errors and 13% fewer treatment errors among clinicians with access to AI Consult. That study compared visits by clinicians with and without the tool. The randomised trial added a follow-up of what happened to patients afterwards, and that is where no clear difference appeared. AI’s best-known win in biology, the AlphaFold Nobel, was for predicting protein structures in a lab, not for treating patients.
What the authors say the study does not show
The Discussion lists the limits itself. The trial was not powered for rare serious harms, and finding no difference does not establish that the two arms were equally safe. There was no prespecified safety framework. Thirty-three serious adverse events occurred, 27 hospitalisations and 6 deaths, and the independent review found none causally linked to the tool.
The results come from a single private, urban network in Nairobi and may not hold in rural or public-sector clinics. Penda’s standards were already high, which may have left less room to improve. The results are tied to one version of GPT-4o, a point the authors describe as a temporal benchmark, and 14 days of follow-up may be too short for downstream effects. Clinicians could not be blinded, and because they shared facilities, informal exchange of ideas between arms was possible. The authors argue that this would push results toward no difference.
The Discussion treats the link from better notes to better patient outcomes as an assumption. The primary outcome did not improve, so the trial did not test that step.
Who paid, and who has a stake
The Gates Foundation funded the trial, with PATH as sponsor. The competing-interests statement says two authors, R.K. and S.K., hold stock options in Penda Health. OpenAI supplied cloud compute credits and technical guidance to Penda for building AI Consult. The statement says the decision to use OpenAI’s product came before that offer, and that OpenAI had no part in the design, data collection, analysis, writing or the decision to publish.
Two readings from outside the trial
NPR quoted two clinicians who were not involved. Dr Jonathan Chen, an associate professor of medicine and biomedical data science at Stanford University, called the study important because it moves beyond simulated tests of AI systems. He said the largest gain may come from timely and consistent access to care, such as help with drafting notes so clinicians can see more patients.
Dr Nicholas Okumu, an orthopaedic surgeon at Kenyatta National Hospital who researches healthcare AI, gave the cautious view. He told NPR, “even AI that’s approved can still cause harm, so oversight has to stay active.”
The Indian context
The Indian Council of Medical Research published its ethical guidelines for AI in biomedical research and healthcare in 2023, a 76-page document from the DHR-ICMR Artificial Intelligence Cell. On 20 February 2026, ICMR and ICMR-NIRDH held a session at Bharat Mandapam, part of the IndiaAI Impact Summit, on the regulatory pathway for AI-based medical devices, bridging training, validation and clinical evaluation. The national programme is covered in India Is Building AI in Its Own Languages.
One Indian trial of a similar kind is registered on ClinicalTrials.gov, and it measures something different. A pilot registered as NCT07432893 ran in Birbhum and Puruliya districts of West Bengal with the Liver Foundation. Each of 736 patients had two consultations in one visit, in random order: one with a nurse using an AI tool, one with a physician. Two blinded physician graders scored each consultation on a clinical reasoning and management rubric on the same day. The record lists the pilot as completed on 1 August 2026. Its primary outcome is the quality of the consultation, not the patient’s health afterwards, and its sponsor is a HEAL India researcher rather than ICMR or AIIMS.
A July 2025 correspondence in Nature Medicine, led by Mateen, observed that only a handful of randomised trials of AI for health had been run in Africa and that none had tested a generative AI tool.
Sources & further reading
- Nature Medicine: Generative AI-enabled clinical decision support system in primary care, a pragmatic, cluster-randomized trial (Agweyu and others, published 26 June 2026)
- Zenodo: Protocol for the multi-facility pragmatic cluster randomized controlled trial in Nairobi, Kenya (Menon and others, June 2025)
- PATH: PATH launches clinical trial on the use of artificial intelligence in primary health care (11 March 2025)
- Nature Medicine: Trials for LLM-supported clinical decisions in African primary healthcare (Mateen and others, correspondence, 3 July 2025)
- arXiv: AI-based Clinical Decision Support for Primary Care, A Real-World Study (Korom and others, 22 July 2025)
- NPR: This AI tool promises a 'second pair of eyes' to clinicians. Did patients benefit? (Joseph Kim, 23 July 2026, as carried by WOSU)
- NPR: This AI tool promises a 'second sight of eyes' to clinicians. Did patients benefit? (KJZZ carriage of the same report)
- News-Medical: Clinical trial evaluates generative AI support tool in primary care (26 June 2026)
- Indian Council of Medical Research: Ethical guidelines for application of Artificial Intelligence in Biomedical Research and Healthcare (2023)
- Indian Council of Medical Research: IndiaAI Impact Summit session on the regulatory pathway for AI-based medical devices (20 February 2026)
- ClinicalTrials.gov: LLM-enabled nurse treatment planning in 2 Indian districts, a pilot study (NCT07432893)
Researched and written with the help of AI tools and edited for accuracy. Provided for general information and discussion only, not professional advice. See our editorial standards and disclaimer. Spotted an error? Tell us.
Enjoyed this? Get the next one.
One good read at a time, straight to your inbox. No spam, unsubscribe anytime.