ARISE
ARISE Logo
Jul 27, 2026Blog

Introducing MAST v1.0: A Framework for Evaluating Medical AI Across Clinical Capabilities

The MAST leaderboard, showing general-purpose AI models ranked by composite clinical score

Today, we are launching the Medical AI Superintelligence Test, or MAST: a living benchmark platform for evaluating medical AI systems on clinically meaningful tasks.

Why MAST?

Clinical competence is not a single capability. A model can perform well on knowledge questions while failing to integrate that knowledge in complex clinical scenarios. It can make the right diagnosis, while recommending an unsafe plan. It can be comprehensive but too aggressive, cautious but incomplete, fluent but brittle, or accurate on average while failing catastrophically in a specific specialty, disease area, imaging modality, or patient subgroup.

Numerous benchmarks — including many from the ARISE network — have sought to assess these various tasks and their underlying model cognitive traits. Yet the deep fragmentation of the landscape has made it difficult to draw durable, reliable conclusions. Leaderboards are not often actively maintained, becoming rapidly out of date given the pace of AI development. Suggestive results on one benchmark are rarely examined in the context of others, rendering it difficult to draw deeper inference.

The goals of MAST are twofold:

First, to draw together a representative set of the field’s highest quality benchmarks into a shared infrastructure, actively maintained and rapidly updated with progress in the broader field.

Second, to enable cross-benchmark inference, allowing for more thorough and comprehensive characterization of model performance and the underlying traits which drive it.

What is MAST?

With expanding model capabilities, the range of tasks that models are being asked to perform has similarly expanded. MAST assesses the full breadth of these capabilities, with benchmarks targeting multiple clinical domains of importance, including:

MAST domains6 domains · 9 benchmarks
DiagnosisDiagnosis cases from the Script Concordance Testing Benchmark (SCT-Diagnostic), as well as the New England Journal of Medicine Clinicopathologic Case Series (CPC-Bench).
Script Concordance Test — DiagnosticProbabilistic diagnostic reasoning under uncertainty — how a single new finding should shift the likelihood of a diagnosis, the way clinicians actually reason with incomplete information.SCT benchmark ↗
CPC-Bench — DiagnosticOpen-ended diagnostic reasoning on real New England Journal of Medicine clinicopathologic case conferences (CPCs) — long, information-dense cases that resolve to a single final diagnosis.NEJM AI paper ↗
The six clinical domains MAST evaluates and the benchmarks behind each. Select a domain to see what its benchmarks measure and where the underlying data comes from.

Many of these benchmarks have undergone further MAST-specific enhancements, including deeper annotation of the nature and content of questions, allowing for more detailed cross-benchmark analyses.

As not every model is designed for every task (in particular, the proliferation of clinical models that lack multimodal capability, or agentic tool use capabilities), we also evaluate and report two MAST subsets:

MAST subsets
MAST-Clinical3 domains
DiagnosisManagementSafety
MAST-General5 domains
DiagnosisManagementSafetyRadiologyMultimodal
The two MAST subsets and the domains feeding each. MAST-Clinical covers diagnosis, management, and safety; MAST-General adds radiology and multimodal reasoning. Hover a subset to trace its domains.

These allow us to compare models fairly against others within their class.

What did we find?

First, the frontier is crowded, but not uniform. The aggregate scores are useful, but incomplete. The highest-performing models are closely clustered on the composite rankings, but the ordering changes across clinical dimensions. There is no single model that dominates every clinically relevant capability.

Score setGPT-5.5OpenAIGPT-5.4OpenAIOpus 4.7AnthropicOpus 4.6AnthropicGemini 3.5 FlashGoogleGemini 3.1 ProGoogleMedGemma 1.5 4BGoogleMedGemma 27BGoogle
Diagnosis75.1%71.9%75.4%75.4%72.1%72.3%45.2%59.1%
Management76.3%70.9%69.2%65.9%56.7%58.3%40.7%47.7%
Safety73.8%73.2%71.1%63.4%63.9%63.9%32.6%46.3%
Multimodal Images42.9%46.5%42.7%40.4%49.6%49.4%16.4%28.7%
Multimodal Radiology57.6%47.8%54.7%45.2%61.7%58.1%57.8%55.3%
Agentic56.7%38.8%40.6%44.4%19.4%10.4%
Column-style MAST score-set matrix. Switch between general-purpose and clinical-model views. Ringed values mark the current point-estimate leader among the models shown. Values are raw benchmark scores, not percentiles.

Second, progress on medical tasks is uneven. Recent AI systems have made dramatic gains on many general benchmarks, but those gains do not translate smoothly into every clinical benchmark. Some domains show clear improvement; others remain flatter, noisier, or more model-specific. This suggests that clinical competence is not a single capability that automatically rises with general model scale.

MAST Clinical ReasoningComposite clinical score0%25%50%75%100%Jan 24Jan 25Jan 26GPT-4 / MAST Clinical Reasoning: 57.1% (2023-05-28)Claude 3 Haiku / MAST Clinical Reasoning: 47.7% (2024-03-13)GPT-4o / MAST Clinical Reasoning: 55.0% (2024-11-20)Gemini 2.0 Flash / MAST Clinical Reasoning: 54.2% (2025-02-05)Llama 4 Maverick / MAST Clinical Reasoning: 53.0% (2025-04-05)Llama 4 Scout / MAST Clinical Reasoning: 49.6% (2025-04-05)GPT-4.1 / MAST Clinical Reasoning: 65.5% (2025-04-14)GPT-4.1 mini / MAST Clinical Reasoning: 62.5% (2025-04-14)Gemini 2.5 Pro / MAST Clinical Reasoning: 57.6% (2025-06-17)Gemini 2.5 Flash / MAST Clinical Reasoning: 57.1% (2025-06-17)Grok 4 / MAST Clinical Reasoning: 64.6% (2025-07-09)MedGemma 27B / MAST Clinical Reasoning: 50.7% (2025-07-09)GPT-5 / MAST Clinical Reasoning: 72.8% (2025-08-07)GPT-5 mini / MAST Clinical Reasoning: 67.0% (2025-08-07)Grok 4 Fast / MAST Clinical Reasoning: 64.7% (2025-09-19)Claude Sonnet 4.5 / MAST Clinical Reasoning: 65.0% (2025-09-29)Claude Haiku 4.5 / MAST Clinical Reasoning: 59.2% (2025-10-01)Opus 4.5 / MAST Clinical Reasoning: 65.3% (2025-11-24)DeepSeek V3.2 / MAST Clinical Reasoning: 55.1% (2025-12-01)GPT-5.2 / MAST Clinical Reasoning: 72.5% (2025-12-10)Gemini 3 Flash / MAST Clinical Reasoning: 59.1% (2025-12-17)MedGemma 1.5 4B / MAST Clinical Reasoning: 40.3% (2026-01-13)Kimi K2.5 / MAST Clinical Reasoning: 64.7% (2026-01-27)Opus 4.6 / MAST Clinical Reasoning: 68.3% (2026-02-04)Claude Sonnet 4.6 / MAST Clinical Reasoning: 67.1% (2026-02-17)Gemini 3.1 Pro / MAST Clinical Reasoning: 63.3% (2026-02-19)GPT-5.4 / MAST Clinical Reasoning: 71.6% (2026-03-05)GPT-5.4 mini / MAST Clinical Reasoning: 64.2% (2026-03-17)Opus 4.7 / MAST Clinical Reasoning: 71.5% (2026-04-16)Kimi K2.6 / MAST Clinical Reasoning: 67.3% (2026-04-20)DeepSeek V4 Pro / MAST Clinical Reasoning: 64.0% (2026-04-23)GPT-5.5 / MAST Clinical Reasoning: 75.5% (2026-04-24)
ARC-AGI-2General abstract reasoning0%25%50%75%100%Jan 24Jan 25Jan 26GPT-4o / ARC-AGI-2: 0.0% (2024-11-20)DeepSeek R1 / ARC-AGI-2: 1.1% (2025-01-20)Gemini 2.0 Flash / ARC-AGI-2: 1.3% (2025-02-05)Claude 3.7 Sonnet / ARC-AGI-2: 0.7% (2025-02-24)Llama 4 Maverick / ARC-AGI-2: 0.0% (2025-04-05)Llama 4 Scout / ARC-AGI-2: 0.0% (2025-04-05)GPT-4.1 / ARC-AGI-2: 0.4% (2025-04-14)GPT-4.1 mini / ARC-AGI-2: 0.0% (2025-04-14)Claude Opus 4 / ARC-AGI-2: 8.6% (2025-05-22)Gemini 2.5 Pro / ARC-AGI-2: 4.9% (2025-06-17)Gemini 2.5 Flash / ARC-AGI-2: 2.5% (2025-06-17)Grok 4 / ARC-AGI-2: 16.0% (2025-07-09)GPT-5 / ARC-AGI-2: 9.9% (2025-08-07)GPT-5 mini / ARC-AGI-2: 4.4% (2025-08-07)Grok 4 Fast / ARC-AGI-2: 5.3% (2025-09-19)Claude Sonnet 4.5 / ARC-AGI-2: 13.6% (2025-09-29)Claude Haiku 4.5 / ARC-AGI-2: 4.0% (2025-10-01)Opus 4.5 / ARC-AGI-2: 37.6% (2025-11-24)DeepSeek V3.2 / ARC-AGI-2: 4.0% (2025-12-01)GPT-5.2 / ARC-AGI-2: 52.9% (2025-12-10)Gemini 3 Flash / ARC-AGI-2: 33.6% (2025-12-17)Kimi K2.5 / ARC-AGI-2: 11.8% (2026-01-27)Opus 4.6 / ARC-AGI-2: 68.8% (2026-02-04)Claude Sonnet 4.6 / ARC-AGI-2: 58.3% (2026-02-17)GPT-5.4 / ARC-AGI-2: 74.0% (2026-03-05)Grok 4.20 / ARC-AGI-2: 65.1% (2026-03-09)GPT-5.4 mini / ARC-AGI-2: 18.9% (2026-03-17)Opus 4.7 / ARC-AGI-2: 75.8% (2026-04-16)GPT-5.5 / ARC-AGI-2: 85.0% (2026-04-24)
MAST Clinical Reasoning has improved more slowly than ARC-AGI-2 over the same model-release window. Select an organisation in the legend to isolate its releases.

Third, models have distinct clinical profiles. A model can be strong at diagnosis and weaker at management; strong in text and weaker in images; strong in static reasoning and weaker in agentic workflows. Small specialist models may outperform much larger generalist systems on narrow modalities, while remaining far less capable across the broader clinical suite.

For example, the tiny 4 billion parameter MedGemma 1.5 4B-IT manages top-tier performance at multimodal radiology, despite demonstrating relatively poor performance in all of our other evaluated domains.

255075100DiagnosisManagementSafetyImagesRadiologyAgentic

5 selected

OpenAI

Other OpenAI

Anthropic

Other Anthropic

Google

Other Google

xAI

Other xAI

Moonshot AI

Other Moonshot AI

DeepSeek

Other DeepSeek

Meta

Other Meta
Percentile-ranked model profiles across MAST domains. Switch between general and clinical model sets, then add or remove lab-grouped models from the radar. Values are percentile ranks within the evaluated field, not raw scores.

What's next

MAST is not a claim that today’s models have reached medical superintelligence. It is a way to measure the path toward it.

Version 1.0 establishes the basic infrastructure: a shared platform, maintained evaluations, standardized model runs, and public reporting across diagnosis, management, safety, radiology, multimodal reasoning, and agentic clinical tasks. That infrastructure matters because one-off benchmark papers age quickly, while models, prompts, scaffolds, and deployment contexts change constantly.

The next phase of MAST will move from benchmark scores to cross-benchmark trait analysis. In clinical practice, two physicians can both be “right” while practicing very differently: one may test broadly and escalate early; another may be more restrained, tolerate uncertainty longer, and avoid low-value intervention. The same is true of AI systems. A model may be accurate but overly aggressive, safe but incomplete, comprehensive but resource-intensive, or highly capable in one specialty while brittle in another.

MAST will increasingly use benchmarks as probes of these underlying clinical behaviors. We are particularly interested in traits such as aggressiveness, restraint, completeness, escalation threshold, diagnostic breadth, resource intensity, calibration, subgroup robustness, and stability across changes in context. These are not fully captured by any single benchmark. They require repeated evaluation across tasks, specialties, modalities, and model families.

That is where MAST is going next: from leaderboards to clinical behavior profiles. We seek to assess not only which model scores highest, but what kind of clinical actor a model is becoming — what it overdoes, what it misses, when it escalates, when it defers, and whether those tendencies are reliable enough for real clinical use.

We welcome model and benchmark submissions that help build that picture.

Acknowledgements

ARISE is an interdisciplinary research network of clinicians, researchers, and builders across academic medical centers. Read more about our team here: https://www.arise-ai.org/team. MAST draws on benchmarks contributed by the teams behind SCT-Bench, CPC-Bench, First Do No Harm, ReXrank, MedAgentBench, and PhysicianBench. A special shout-out to our members of technical staff, whose engineering work built and maintains the MAST platform.