Research

Explore our publications and preprints advancing healthcare through rigorous AI evaluation.

Preprint
Jul 7, 2026

SycoEval-EM: Sycophancy Evaluation of Large Language Models in Simulated Clinical Encounters for Emergency Care

Large language models (LLMs) deployed in clinical decision support risk acquiescing to patient pressure for guideline-discordant care. We developed SycoEval-EM, […]

Preprint
May 24, 2026

Teaching large language models to reason like expert diagnosticians

Differential diagnosis is an iterative process that integrates patient information with broader medical knowledge. Clinical case series such as the […]

Preprint
Mar 31, 2026

Towards Conversational Medical AI with Eyes, Ears and a Voice

The practice of medicine relies not only upon skillful dialogue but also on the nuanced exchangeand interpretation of rich auditory […]

Preprint
Mar 18, 2026

Deployment and Evaluation of an EHR-integrated, Large Language Model-Powered Tool to Triage Surgical Patients

Surgical co-management (SCM) is an evidence-based model in which hospitalists jointly manage medically complex perioperative patients alongside surgical teams. Despite […]

Preprint
Mar 14, 2026

Clinician input steers frontier AI models toward both accurate and harmful decisions

Large language models (LLMs) are entering clinician workflows, yet evaluations rarely measure how clinician reasoning shapes model behavior during clinical […]

Preprint
Feb 6, 2026

MedAgentBrief for Hospital Course Summarization: Safety, Use, and Discharge Documentation Burden

Importance: High-quality discharge summaries are essential for safe care transitions but contribute substantially to clinician documentation burden and burnout. While […]

Preprint
Jan 21, 2026

Adoption and Use of LLMs at an Academic Medical Center

While large language models (LLMs) can support clinical documentation needs, standalone tools struggle with “workflow friction” from manual data entry. […]

Preprint
Jan 10, 2026

MedAgentBench v2: Improving Medical LLM Agent Design

MedAgentBench is the first benchmark for evaluating LLM agents on clinical tasks in a FHIR-compliant EHR. In this paper, we […]

Preprint
Dec 4, 2025

SmartAlert: Implementing Machine Learning-Driven Clinical Decision Support for Inpatient Lab Utilization Reduction

Repetitive laboratory testing unlikely to yield clinically useful information is a common practice that burdens patients and increases healthcare costs. […]

Preprint
Dec 1, 2025

First, do NOHARM: towards clinically safe large language models

Large language models (LLMs) are routinely used by physicians and patients for medical advice, yet their clinical safety profiles remain […]

Latest News

View all

Get the latest on our studies, grant awards, and media coverage.