Evidence OSMedical evidence & research
Evidence OSFrontiers
FRONTIERS

Frontiers

Follow important advances in medical evidence, clinical AI and health decisions. Understand the findings, their relevance to patients and clinicians, and what still needs evaluation.

Daily brief · compiled
npj Breast Cancer | Retrospective clinical-trial data study

AI-extracted health data should remain traceable to the source

The study compared manual entry with AI extraction of breast-cancer trial data from EHRs. Each extracted variable retained a source-text location for reviewer verification.

What it means for health and care

For clinicians and research clients, source-linked results make misreadings easier to detect, disagreements easier to resolve and corrections easier to record—an important safety condition beyond efficiency.

Study limits and provenance

This was a retrospective comparison in one hospital and one cancer type. It does not establish reliable performance across hospitals, lower costs or better patient outcomes.

Artificial intelligence for clinical data extraction from EHRs in breast cancer trials: the MIRROR study
npj Artificial Intelligence | Adversarial benchmark study

Medical AI can be steered even when it cites evidence

The study shows that adding a few crafted documents to a retrieval corpus can change how medical AI frames the same facts and steer answers toward a product or viewpoint.

What it means for health and care

Patients and clinicians need more than the presence of citations: who produced the sources, whether independent sources agree, whether commercial framing is present and whether conclusions change when sources are replaced.

Study limits and provenance

This is a laboratory benchmark, not evidence that the same attack has occurred in clinical practice. Real-world frequency, patient harm and the best defence remain unproven.

FramePoison: attacks on medical RAG that target framing, not just facts
npj Digital Surgery | Scoping review

Very little published clinical evidence supports AI used during surgery

Researchers screened 3,020 records and found only five studies of real intraoperative decision support. Only one had completed results, based on five patients.

What it means for health and care

Patients, surgeons and hospitals should not treat technical demonstrations, registered trials or clinician final authority as proof of safety or benefit. Auditability, accountability and outcome evaluation remain necessary.

Study limits and provenance

The review was limited to intraoperative AI, included few heterogeneous studies, and proposed an externally unvalidated governance scorecard.

Ethical considerations for intraoperative implementation of artificial intelligence clinical decision support systems: a scoping review
npj Digital Medicine | Exploratory multicentre randomised trial

AI-assisted telerehabilitation did not improve the primary motor outcome

One hundred and twenty people with early Parkinson disease were randomised to different training durations, with 71 in the final analysis. The primary motor score did not differ between groups at three months; no falls or other adverse events were recorded.

What it means for health and care

The negative result reminds patients and clinicians that using AI—even within a randomised trial—does not establish benefit. Prespecified primary outcomes and appropriate controls matter more than post-hoc findings.

Study limits and provenance

Attrition was substantial, with no usual-care or non-AI control. Comparing training durations does not isolate the added value of AI.

AI-assisted telerehabilitation in early Parkinson’s disease: a multicenter, randomized, multi-arm comparative trial
Daily brief · compiled
UK Government and MHRA · Healthcare AI regulation update

Healthcare AI will need continuing evidence after approval

On 6 October, the UK Government accepted all 44 healthcare AI regulation recommendations and opened a new AI Airlock phase focused on post-market and lifecycle monitoring.

What it means for health and care

Patients, clinicians and buyers will need to know not only whether a system was authorized, but whether updates are revalidated, real-world use is continuously monitored and problems are corrected.

Study limits and provenance

This is a policy and regulatory-sandbox programme, not completed legislation or evidence of improved patient outcomes.

Government backs recommendations of NHS doctors-led AI Commission
npj Digital Medicine · Randomized clinical trial

AI can collect a fuller history, while examination and synthesis remain clinical responsibilities

In a randomized trial of 172 ophthalmology patients, AI achieved more complete histories and higher patience and empathy ratings but took longer; clinicians gained more when ocular examination findings were added.

What it means for health and care

A more appropriate workflow may use AI for standardized information collection followed by clinician examination, synthesis and accountability, rather than autonomous end-to-end care.

Study limits and provenance

The study came from one ophthalmology hospital, was small, and did not assess long-term outcomes or real-world clinical safety.

Standardized pre-consultation by a large language model agent vs ophthalmology residents: a randomized clinical trial
medRxiv · Non-peer-reviewed real-world research manuscript

Clinician edits to AI notes can help reveal system changes

A study of 268,379 outpatient encounters found that clinical-concept edits can identify meaningful changes at scale and monitor shifts around system rollout.

What it means for health and care

For hospitals, recording what clinicians changed and why may reveal quality changes more effectively than measuring time savings alone.

Study limits and provenance

The study is not peer reviewed and came from one institution; edits are not equivalent to errors and have not been shown to reduce patient harm.

A Framework to Monitor Editing of Artificial Intelligence-Generated Medical Documentation
npj Digital Medicine · Systematic evidence survey

Authorized medical AI may still lack publicly checkable performance evidence

Among 77 CE-marked or FDA-authorized pathology and hematology AI devices, only 39% had identifiable public device-specific performance evidence.

What it means for health and care

Hospitals and customers should look beyond authorization to public performance data, intended populations, evaluation metrics and local validation.

Study limits and provenance

Failure to find public evidence does not mean regulators reviewed no data and cannot establish that a product is ineffective or unsafe.

Availability of performance evidence of approved AI diagnostic software in pathology and hematology morphology
Daily brief · compiled
JMIR · Systematic review and Bayesian network meta-analysis of randomised trials

More complex hospital alerts do not guarantee patient benefit

A review published on 24 September included 28 randomised trials. Within the strict digital-surveillance evidence, rule-based electronic monitoring, predictive models and continuous physiological monitoring did not clearly reduce death or ICU transfer.

What it means for health and care

Hospitals, clinicians and patients should look beyond whether a model detects risk to who receives the alert, whether action follows promptly, and whether important outcomes improve.

Study limits and provenance

The evidence networks were sparse, lacked head-to-head active comparisons, and credible intervals included both benefit and harm; overall confidence was very low.

Clinical Surveillance Technologies in Nonintensive Care Unit Hospital Settings: Systematic Review and Bayesian Network Meta-Analysis of Randomized Trials
Hospital announcement · Individualised clinical AI

Connecting kidney-risk prediction with individual follow-up

Hospital Clínic Barcelona described Renal-Trust on 5 October, combining secure clinical data infrastructure with explainable kidney-progression prediction using information from more than 4,000 patients. The award event occurred on 1 October.

What it means for health and care

The potential value is earlier identification of people who may need closer follow-up and clearer explanations of risk; real-workflow evaluation is still needed.

Study limits and provenance

The institutional announcement provides no prospective comparative trial establishing improved patient outcomes.

医院公告 · 个体化临床 AI
Frontiers · Systematic review

Antibiotic decision AI still needs stronger clinical evidence

A 30 September systematic review included ten studies, with clinical deployment in only three. Some favourable signals were reported, but outcome certainty was low or very low and heterogeneity prevented pooling.

What it means for health and care

Clinicians and hospital buyers should distinguish model performance, real-workflow use and patient outcomes, including the infections and populations covered.

Study limits and provenance

The review does not establish effectiveness across antibiotic AI systems; prospective multicentre evaluation remains needed.

Frontiers · 系统综述
JMIR · Retrospective development and validation

Prioritising serious events in large patient-safety reporting systems

A 29 September study used 101,239 retrospective safety reports from a Canadian academic health system. Text-based ranking models better prioritised high-severity events for investigation.

What it means for health and care

The potential value is focusing limited safety-investigation resources on consequential reports; this still needs validation in practice.

Study limits and provenance

Retrospective ranking performance does not demonstrate fewer patient harms; prospective workflow and external evaluations are needed.

JMIR · 回顾性开发验证
Daily brief · compiled
npj Digital Medicine · Perspective

How can a clinical AI recommendation be checked?

A medical perspective argues that claims informing care need adequate evidence, precisely verifiable sources and presentation suited to reviewers such as patients and clinicians.

What it means for health and care

Patients can ask which study supports advice and whether it applies to them. Clinicians need to locate the supporting passages and limits. Clear evidence trails support discussion of reasoning, risks and uncertainty.

Study limits and provenance

Published 19 September 2026. A perspective proposing a verification framework, without clinical evidence of patient benefit.

Toward reviewable medical evidence synthesis for care delivery
Scientific Reports · Benchmark study

How can medical AI look for missing evidence?

The MRER multi-agent framework uses retrieved evidence to identify unresolved questions and guide further searches. The study reported mean accuracy of 70.68% across three medical question-answering benchmarks.

What it means for health and care

For complex health questions, patients and clinicians need to know what has support, where evidence is missing and what further searches add. Explicit gaps clarify the need for further verification or professional review.

Study limits and provenance

Published 19 September 2026, accepted 16 September; an early publisher version. Question-answering benchmarks do not establish clinical benefit.

Retrieval-augmented multi-agent framework for evidence-centric medical reasoning
Nature Medicine · Diagnostic benchmark study

Beyond accuracy, how many cases can AI answer?

An on-premise clinical-agent study retained 49.4% of cases using a consistency threshold. Diagnostic accuracy was 98.9% within that subset. Both figures need to be read together.

What it means for health and care

Evaluate accuracy alongside population coverage and when professional review is needed. Clear scope and referral conditions support appropriate use and prevent selected-case performance from being interpreted as performance for all cases.

Study limits and provenance

Published 15 September 2026. Simulation and diagnostic benchmarks do not establish routine clinical effectiveness or improved patient outcomes.

On-premise medical AI agents for reliable clinical decision-making