Claude Mythos Preview Claude Opus 4.6 Claude Sonnet 4.6 GPT-5.4 Gemini 3.1 Pro Expert Human Baseline

Overall Results Summary

Overview

Claude Mythos Preview compared against the second-best model on each benchmark.

Software Engineering: SWE-bench Series

Coding

Measures an AI's ability to solve real-world software engineering problems. Tasks come from genuine bugs in open-source repositories; the model must generate code patches that pass the test suite.

Verified
500 problems verified by human engineers as solvable, curated by OpenAI. The most widely used software engineering benchmark.
93.9%
Pro
Problems from actively maintained repositories involving larger multi-file diffs with no public answer leakage — the truest test of real-world engineering capability.
77.8%
Multilingual
Extends the format to 300 problems across 9 programming languages (Python, Java, TypeScript, and more), testing cross-language generalization.
87.3%
Multimodal
Adds screenshots and design mockups to issue descriptions, testing the model's ability to complete engineering tasks by integrating visual and textual context.
59.0%

Mythos Preview leads all prior models by a wide margin on every variant. After removing suspected memorized problems, the lead narrows by at most 3.5 pp and rankings are unchanged.

Terminal Operations: Terminal-Bench 2.0

Agentic

Developed by Stanford University and the Laude Institute. Tests an AI's ability to complete real tasks in terminal and command-line environments, covering 89 system-operation scenarios, each run inside an isolated Kubernetes container. The benchmark is sensitive to inference speed.

With the 2.1 patch applied and the timeout extended to 4 hours, Mythos Preview jumped from 82% to 92.1%, widening the gap over GPT-5.4 (75.3%) even further.

Scientific Reasoning: GPQA Diamond

Reasoning

"Google-proof" graduate-level science multiple-choice questions — 198 problems spanning physics, chemistry, and biology. Domain experts can answer them correctly, but non-experts find them very difficult even with search engine access.

The four top models are tightly clustered (gap ≤ 3.2 pp), indicating that GPQA Diamond is approaching saturation as a discriminator for today's strongest models.

Multilingual Knowledge: MMMLU

Knowledge

A comprehensive knowledge and reasoning test spanning 57 subjects across 14 non-English languages, evaluating breadth of knowledge and consistency of reasoning beyond English.

Mathematical Proof: USAMO 2026

Math

The top high-school mathematics competition in the United States: 6 open-ended proof problems. The exam was held March 21–22, 2026 — after model training cutoffs — ruling out data contamination. A judging panel composed of three frontier models was used, with the lowest score taken as the final result.

The largest gap in the evaluation: Opus 4.6 scored only 42.3%, while Mythos Preview reached 97.6% — a generational leap in mathematical reasoning. Gemini 3.1 Pro independently confirmed scoring on 58 of 60 problems.

Long-Context Reasoning: GraphWalks

Long Context

Tests multi-hop reasoning ability under extremely long contexts (256K–1M tokens). The context contains a directed graph; the model must perform breadth-first search (BFS) or find parent nodes through pure reasoning with no shortcuts.

On the BFS task, Mythos Preview scores more than twice Opus 4.6 and nearly four times GPT-5.4. Results cannot be reproduced via public APIs (exceeds the 1M-token limit).

Agentic Search: Humanity's Last Exam

HLE

2,500 multimodal questions at the frontier of human knowledge, billed as "the hardest AI benchmark." Evaluated in two modes: reasoning-only (no tools) and agentic (with web search, code execution, and other tools).

Providing tools improves scores by ~8 pp; Mythos Preview's advantage over competitors is even clearer in the tool-equipped mode, demonstrating superior coordination of tool use and reasoning.

Deep Information Retrieval: BrowseComp

Search

Tests an agent's ability to locate "extremely hard-to-find information" on the open web, evaluating multi-round search, page scraping, and reasoning coordination rather than stored knowledge.

BrowseComp test-time compute scaling

Figure 6.10.2.A — Accuracy increases monotonically with token budget. Mythos Preview achieves a higher accuracy using 226K tokens — just 1/4.9 of what Opus 4.6 requires.

Multimodal: LAB-Bench FigQA

Multimodal

Developed by FutureHouse. Tests whether models can correctly interpret complex scientific figures from biology research papers. Part of the LAB-Bench series evaluating AI assistance for real scientific research.

Mythos Preview's tool-free score (79.7%) already surpasses the expert human baseline (77.0%); with Python image-processing tools it leaps to 89.0%.

Multimodal: ScreenSpot-Pro

GUI

Developed by the National University of Singapore. Tests whether models can precisely locate UI elements in high-resolution screenshots of professional desktop software, covering 23 specialized applications. Target elements occupy on average less than 0.1% of the screen area.

Adding Python tools lifts the score from 79.5% to 92.8%, highlighting the critical role of image-processing capability for high-precision GUI grounding.

Multimodal: CharXiv Reasoning

Multimodal

2,323 real figures from arXiv papers across 8 scientific domains. Tests whether models can integrate visual information for multi-step reasoning — the closest existing multimodal benchmark to real scientific reading.

Mythos Preview without tools (86.1%) outperforms Opus 4.6 with tools (78.9%) by 7.2 pp — a striking generational gap.

Computer Automation: OSWorld

Computer Use

Agents operate inside a real Ubuntu virtual machine, using mouse and keyboard to complete actual computer tasks: editing documents, browsing the web, managing files. Runs at 1080p resolution with a maximum of 100 actions per task.

Data Notes
· All Mythos Preview scores: adaptive thinking, max effort, default sampling settings, mean of 5 independent trials
· Competitor data sourced from each model's official System Card or benchmark leaderboard
· In Terminal-Bench 2.0, OpenAI used a proprietary harness; comparisons with other models are for reference only
· GraphWalks results cannot be reproduced via public APIs (exceeds the 1M-token limit)
· "—" indicates the model did not report a result on that benchmark