Explore our publications and preprints advancing healthcare through rigorous AI evaluation.
Doctors increasingly rely on AI in the clinic, yet which report features make AI-generated responses useful and trustworthy remains unclear. […]
Medical artificial intelligence (AI) is increasingly developed, piloted and used in clinical practice, yet the translation of technical capability into […]
Surgical comanagement (SCM) is an evidence-based care model in which hospitalists jointly manage medically complex perioperative patients alongside surgical teams. […]
We present HealMed, an expert-reviewed benchmark for multilingual evaluation of large language models in medicine. HealMed contains 1,000 examples in […]
Advances in large language models (LLMs) have accelerated medical benchmarking, yet most evaluation of LLMs still relies on exam-style question […]
Automated radiology report generation is advancing rapidly in response to the shortage of radiologists, yet unlike a perception model, existing […]
With the growing use of language models (LMs) in clinical environments, there is an immediate need to evaluate the accuracy […]
High-stakes decisions under uncertainty, such as medical emergency triage, require more than accurate predictions. They depend on estimating the likelihood […]
The development of two medical AI assistants highlights an unnerving challenge: as the technology races ahead, what is the best […]
Researchers urgently need a rigorous, task-based framework to define and measure medical AI ‘superintelligence’, because existing benchmarks are misleading and […]
Get the latest on our studies, grant awards, and media coverage.