Explore our publications and preprints advancing healthcare through rigorous AI evaluation.
Large language models (LLMs) deployed in clinical decision support risk acquiescing to patient pressure for guideline-discordant care. We developed SycoEval-EM, […]
Differential diagnosis is an iterative process that integrates patient information with broader medical knowledge. Clinical case series such as the […]
The practice of medicine relies not only upon skillful dialogue but also on the nuanced exchangeand interpretation of rich auditory […]
Surgical co-management (SCM) is an evidence-based model in which hospitalists jointly manage medically complex perioperative patients alongside surgical teams. Despite […]
Large language models (LLMs) are entering clinician workflows, yet evaluations rarely measure how clinician reasoning shapes model behavior during clinical […]
Importance: High-quality discharge summaries are essential for safe care transitions but contribute substantially to clinician documentation burden and burnout. While […]
While large language models (LLMs) can support clinical documentation needs, standalone tools struggle with “workflow friction” from manual data entry. […]
MedAgentBench is the first benchmark for evaluating LLM agents on clinical tasks in a FHIR-compliant EHR. In this paper, we […]
Repetitive laboratory testing unlikely to yield clinically useful information is a common practice that burdens patients and increases healthcare costs. […]
Large language models (LLMs) are routinely used by physicians and patients for medical advice, yet their clinical safety profiles remain […]
Get the latest on our studies, grant awards, and media coverage.