Explore our publications and preprints advancing healthcare through rigorous AI evaluation.
Large language models (LLMs) deployed in clinical decision support risk acquiescing to patient pressure for guideline-discordant care. We developed SycoEval-EM, […]
Emergency department (ED) quality review often uses administrative electronic triggers (eTriggers), but yields on detecting missed opportunities for diagnosis (MODs) […]
Large frontier models such as GPT-5 and Gemini have demonstrated remarkable performance in a wide range of health application benchmarks. […]
Conventional artificial intelligence has achieved remarkable feats in identifying associations and predictive patterns, yet its limitations in causal reasoning present […]
While large language models (LLMs) have shown promise in diagnostic dialogue1, their capabilities for effective management reasoning—including disease progression, therapeutic […]
Large language models (LLMs) are evolving rapidly and hold great promise for medical applications, yet benchmarking on real-world clinical data […]
Much of modern healthcare remains reactive: patients present with symptoms and clinicians respond. What if, instead, we could anticipate illnesses […]
Multimodal foundation models have shown compelling but conflicting performance in medical image interpretation. However, the ways in which these models […]
As capabilities of artificial intelligence (AI) advance rapidly, human understanding of these systems is increasingly falling behind. Several trends are […]
Differential diagnosis is an iterative process that integrates patient information with broader medical knowledge. Clinical case series such as the […]
Get the latest on our studies, grant awards, and media coverage.