Introducing MAST v1.0: A Framework for Evaluating Medical AI Across Clinical Capabilities

Today, we are launching the Medical AI Superintelligence Test, or MAST: a living benchmark platform for evaluating medical AI systems on clinically meaningful tasks.
Why MAST?
Clinical competence is not a single capability. A model can perform well on knowledge questions while failing to integrate that knowledge in complex clinical scenarios. It can make the right diagnosis, while recommending an unsafe plan. It can be comprehensive but too aggressive, cautious but incomplete, fluent but brittle, or accurate on average while failing catastrophically in a specific specialty, disease area, imaging modality, or patient subgroup.
Numerous benchmarks — including many from the ARISE network — have sought to assess these various tasks and their underlying model cognitive traits. Yet the deep fragmentation of the landscape has made it difficult to draw durable, reliable conclusions. Leaderboards are not often actively maintained, becoming rapidly out of date given the pace of AI development. Suggestive results on one benchmark are rarely examined in the context of others, rendering it difficult to draw deeper inference.
The goals of MAST are twofold:
First, to draw together a representative set of the field’s highest quality benchmarks into a shared infrastructure, actively maintained and rapidly updated with progress in the broader field.
Second, to enable cross-benchmark inference, allowing for more thorough and comprehensive characterization of model performance and the underlying traits which drive it.
What is MAST?
With expanding model capabilities, the range of tasks that models are being asked to perform has similarly expanded. MAST assesses the full breadth of these capabilities, with benchmarks targeting multiple clinical domains of importance, including:
Many of these benchmarks have undergone further MAST-specific enhancements, including deeper annotation of the nature and content of questions, allowing for more detailed cross-benchmark analyses.
As not every model is designed for every task (in particular, the proliferation of clinical models that lack multimodal capability, or agentic tool use capabilities), we also evaluate and report two MAST subsets:
These allow us to compare models fairly against others within their class.
What did we find?
First, the frontier is crowded, but not uniform. The aggregate scores are useful, but incomplete. The highest-performing models are closely clustered on the composite rankings, but the ordering changes across clinical dimensions. There is no single model that dominates every clinically relevant capability.
| Score set | GPT-5.5OpenAI | GPT-5.4OpenAI | Opus 4.7Anthropic | Opus 4.6Anthropic | Gemini 3.5 FlashGoogle | Gemini 3.1 ProGoogle | MedGemma 1.5 4BGoogle | MedGemma 27BGoogle |
|---|---|---|---|---|---|---|---|---|
| Diagnosis | 75.1% | 71.9% | 75.4% | 75.4% | 72.1% | 72.3% | 45.2% | 59.1% |
| Management | 76.3% | 70.9% | 69.2% | 65.9% | 56.7% | 58.3% | 40.7% | 47.7% |
| Safety | 73.8% | 73.2% | 71.1% | 63.4% | 63.9% | 63.9% | 32.6% | 46.3% |
| Multimodal Images | 42.9% | 46.5% | 42.7% | 40.4% | 49.6% | 49.4% | 16.4% | 28.7% |
| Multimodal Radiology | 57.6% | 47.8% | 54.7% | 45.2% | 61.7% | 58.1% | 57.8% | 55.3% |
| Agentic | 56.7% | 38.8% | 40.6% | 44.4% | 19.4% | 10.4% | — | — |
Second, progress on medical tasks is uneven. Recent AI systems have made dramatic gains on many general benchmarks, but those gains do not translate smoothly into every clinical benchmark. Some domains show clear improvement; others remain flatter, noisier, or more model-specific. This suggests that clinical competence is not a single capability that automatically rises with general model scale.
Third, models have distinct clinical profiles. A model can be strong at diagnosis and weaker at management; strong in text and weaker in images; strong in static reasoning and weaker in agentic workflows. Small specialist models may outperform much larger generalist systems on narrow modalities, while remaining far less capable across the broader clinical suite.
For example, the tiny 4 billion parameter MedGemma 1.5 4B-IT manages top-tier performance at multimodal radiology, despite demonstrating relatively poor performance in all of our other evaluated domains.
5 selected
OpenAI
Other OpenAI
Anthropic
Other Anthropic
Other Google
xAI
Other xAI
Moonshot AI
Other Moonshot AI
DeepSeek
Other DeepSeek
Meta
Other Meta
What's next
MAST is not a claim that today’s models have reached medical superintelligence. It is a way to measure the path toward it.
Version 1.0 establishes the basic infrastructure: a shared platform, maintained evaluations, standardized model runs, and public reporting across diagnosis, management, safety, radiology, multimodal reasoning, and agentic clinical tasks. That infrastructure matters because one-off benchmark papers age quickly, while models, prompts, scaffolds, and deployment contexts change constantly.
The next phase of MAST will move from benchmark scores to cross-benchmark trait analysis. In clinical practice, two physicians can both be “right” while practicing very differently: one may test broadly and escalate early; another may be more restrained, tolerate uncertainty longer, and avoid low-value intervention. The same is true of AI systems. A model may be accurate but overly aggressive, safe but incomplete, comprehensive but resource-intensive, or highly capable in one specialty while brittle in another.
MAST will increasingly use benchmarks as probes of these underlying clinical behaviors. We are particularly interested in traits such as aggressiveness, restraint, completeness, escalation threshold, diagnostic breadth, resource intensity, calibration, subgroup robustness, and stability across changes in context. These are not fully captured by any single benchmark. They require repeated evaluation across tasks, specialties, modalities, and model families.
That is where MAST is going next: from leaderboards to clinical behavior profiles. We seek to assess not only which model scores highest, but what kind of clinical actor a model is becoming — what it overdoes, what it misses, when it escalates, when it defers, and whether those tendencies are reliable enough for real clinical use.
We welcome model and benchmark submissions that help build that picture.
Acknowledgements
ARISE is an interdisciplinary research network of clinicians, researchers, and builders across academic medical centers. Read more about our team here: https://www.arise-ai.org/team. MAST draws on benchmarks contributed by the teams behind SCT-Bench, CPC-Bench, First Do No Harm, ReXrank, MedAgentBench, and PhysicianBench. A special shout-out to our members of technical staff, whose engineering work built and maintains the MAST platform.


