NEJM AI

@ai.nejm.org

NEJM AI, a new monthly journal from the publisher of @nejm.org, explores the cutting-edge applications of artificial intelligence and machine learning in clinical medicine. Online at ai.nejm.org.

In a new study on value-informed proxy decision support, Nolan and colleagues examine whether LLMs can improve alignment with patient preferences, extract values from clinical notes, and maintain performance across different reading levels. Full results: https://nejm.ai/3UfoOc9

A bar chart showing the average large language model–patient agreement (blue bars) compared with average human proxy–patient agreement (horizontal brown and black dotted lines) across 10 repetition trials (±95% confidence intervals).

A new Perspective proposes a five-phase framework for evaluating the evidentiary maturity of clinical AI to help distinguish technical performance from clinical readiness and align claims with evidence needed to support safe, effective use. Learn more: https://nejm.ai/4yIi4CK

A table showing phase-based framework for evidentiary maturity in clinical artificial intelligence. The table includes five phases: Model development and representation learning, Internal validation, External validation, Decision impact and implementation readiness, and Postdeployment monitoring. Each phase details core questions, evidence expected, and example claims supported, from technical feasibility to durable clinical effectiveness.

A new Policy Corner analyzes how CMS’s repeal of the NTAP alternative pathway may exacerbate an existing imbalance in the U.S. market by stifling competition and driving health systems toward inferior but built-in EHR vendor–developed AI. Learn more: https://nejm.ai/4r7fjbd

A figure showing the New Technology Add-On Payment (NTAP) policy timeline and approvals by fiscal year. The annual number of NTAP approvals through the traditional (blue) and alternative (reddish brown) pathways is graphed alongside a timeline of major NTAP-related policy events. The number of NTAP technologies approved through the alternative pathway has steadily increased since the pathway was first established in 2020.

Drugs. Devices. And a third modality: AI-based interventions. On NEJM AI Grand Rounds, Dr. Suchi Saria envisions software protocols that identify precisely when to act, what to do, and which patient needs it — while making delivery easier at scale. Hear more from Dr. Saria: https://nejm.ai/ep45

A new study examines parents’ perceptions of a secure LLM chatbot for pediatric cancer and vascular anomalies, highlighting opportunities for caregiver support alongside challenges related to trust, overreliance, emotional readiness, and implementation. Learn more: https://nejm.ai/4zChqYG

A table showing parents’ recommendations for chatbot developers. It includes two columns: "Theme" and "Recommendations." The themes listed are "Tailored information," "Response format," "Queries," and "Chat history." Recommendations include ensuring accuracy, providing simplified information, offering summaries, following up responses, allowing search within the chatbot, and letting users delete chat histories.

A new study analyzing the impact of clinicians modifying AI-generated drafts within an electronic health record showed that complex edits require the most additional time while common administrative edits contribute the greatest workload. See how: https://nejm.ai/4c7M3Lv

A table showing the cumulative clinician time spent editing responses by modification category, April 2024–August 2025. Categories include scheduling, lifestyle advice, and laboratory results, among others. The table details the percentage increase per message, number of messages, and cumulative increase in response time (hours per 1000 physicians).

Documentation is not the holy grail. Better care is. On NEJM AI Grand Rounds, Dr. Suchi Saria argues that clinical AI must move into the center of the encounter, where rigorous tools can help clinicians recognize risk and act sooner. Listen to the full episode: https://nejm.ai/ep45

A new study evaluates a value-informed LLM framework designed to support surrogate decision-making for patients who lose decisional capacity by comparing LLM-generated treatment recommendations with choices made by patients and their proxies. Learn more: https://nejm.ai/3UfoOc9

A bar chart showing the average large language model–patient agreement (blue bars) compared with average human proxy–patient agreement (horizontal brown and black dotted lines) across 10 repetition trials (±95% confidence intervals).

On NEJM AI Grand Rounds, Dr. Suchi Saria explains why health systems need evidence that AI can improve outcomes, reduce utilization, and support a sustainable path to implementation. Listen to the full episode: https://nejm.ai/ep45

Current medical AI frameworks do not sufficiently address patient safety because harms are often delayed, difficult to attribute, and irreversible. A new Perspective proposes a four-tier classification of medical AI safety. Learn more: https://nejm.ai/4wo9oQ4

The image features a quote from an NEJM AI Perspective by B. Sheng et al.: "Medical AI should be judged not only by its performance or safety today, but by whether health care systems can detect, attribute, and mitigate the harms it may cause tomorrow." A colorful silhouette of a face with horizontal lines is on the left. The NEJM AI logo is at the bottom right.

A new Perspective examines how agentic AI may transform biomedical research by shifting bottlenecks from analysis itself to the data infrastructure, governance, and institutional capabilities required to deploy AI effectively. Learn more: https://nejm.ai/4xIiMiR

This image is a flowchart titled "Worked Example of an Agentic Research Workflow from Clinical Question to Auditable Output." A clinician-scientist’s question — illustrated here by a multiple sclerosis example — can be translated into a multistep agentic workflow that defines a computable cohort, searches a secure data environment, maps concepts to available data, retrieves external knowledge, drafts an analysis plan, generates code for supervised execution, produces an auditable report, and routes results to human review.

What changes when an abstract research problem becomes a family tragedy? On NEJM AI Grand Rounds, Dr. Suchi Saria explains how losing her nephew to sepsis led her to build Bayesian and pursue earlier, actionable recognition at the point of care. Listen to the full episode: https://nejm.ai/ep45