Overall Results Summary
OverviewClaude Mythos Preview compared against the second-best model on each benchmark.
Software Engineering: SWE-bench Series
CodingMeasures an AI's ability to solve real-world software engineering problems. Tasks come from genuine bugs in open-source repositories; the model must generate code patches that pass the test suite.
Mythos Preview leads all prior models by a wide margin on every variant. After removing suspected memorized problems, the lead narrows by at most 3.5 pp and rankings are unchanged.
Terminal Operations: Terminal-Bench 2.0
AgenticDeveloped by Stanford University and the Laude Institute. Tests an AI's ability to complete real tasks in terminal and command-line environments, covering 89 system-operation scenarios, each run inside an isolated Kubernetes container. The benchmark is sensitive to inference speed.
With the 2.1 patch applied and the timeout extended to 4 hours, Mythos Preview jumped from 82% to 92.1%, widening the gap over GPT-5.4 (75.3%) even further.
Scientific Reasoning: GPQA Diamond
Reasoning"Google-proof" graduate-level science multiple-choice questions — 198 problems spanning physics, chemistry, and biology. Domain experts can answer them correctly, but non-experts find them very difficult even with search engine access.
The four top models are tightly clustered (gap ≤ 3.2 pp), indicating that GPQA Diamond is approaching saturation as a discriminator for today's strongest models.
Multilingual Knowledge: MMMLU
KnowledgeA comprehensive knowledge and reasoning test spanning 57 subjects across 14 non-English languages, evaluating breadth of knowledge and consistency of reasoning beyond English.
Mathematical Proof: USAMO 2026
MathThe top high-school mathematics competition in the United States: 6 open-ended proof problems. The exam was held March 21–22, 2026 — after model training cutoffs — ruling out data contamination. A judging panel composed of three frontier models was used, with the lowest score taken as the final result.
The largest gap in the evaluation: Opus 4.6 scored only 42.3%, while Mythos Preview reached 97.6% — a generational leap in mathematical reasoning. Gemini 3.1 Pro independently confirmed scoring on 58 of 60 problems.
Long-Context Reasoning: GraphWalks
Long ContextTests multi-hop reasoning ability under extremely long contexts (256K–1M tokens). The context contains a directed graph; the model must perform breadth-first search (BFS) or find parent nodes through pure reasoning with no shortcuts.
On the BFS task, Mythos Preview scores more than twice Opus 4.6 and nearly four times GPT-5.4. Results cannot be reproduced via public APIs (exceeds the 1M-token limit).
Agentic Search: Humanity's Last Exam
HLE2,500 multimodal questions at the frontier of human knowledge, billed as "the hardest AI benchmark." Evaluated in two modes: reasoning-only (no tools) and agentic (with web search, code execution, and other tools).
Providing tools improves scores by ~8 pp; Mythos Preview's advantage over competitors is even clearer in the tool-equipped mode, demonstrating superior coordination of tool use and reasoning.
Deep Information Retrieval: BrowseComp
SearchTests an agent's ability to locate "extremely hard-to-find information" on the open web, evaluating multi-round search, page scraping, and reasoning coordination rather than stored knowledge.
Figure 6.10.2.A — Accuracy increases monotonically with token budget. Mythos Preview achieves a higher accuracy using 226K tokens — just 1/4.9 of what Opus 4.6 requires.
Multimodal: LAB-Bench FigQA
MultimodalDeveloped by FutureHouse. Tests whether models can correctly interpret complex scientific figures from biology research papers. Part of the LAB-Bench series evaluating AI assistance for real scientific research.
Mythos Preview's tool-free score (79.7%) already surpasses the expert human baseline (77.0%); with Python image-processing tools it leaps to 89.0%.
Multimodal: ScreenSpot-Pro
GUIDeveloped by the National University of Singapore. Tests whether models can precisely locate UI elements in high-resolution screenshots of professional desktop software, covering 23 specialized applications. Target elements occupy on average less than 0.1% of the screen area.
Adding Python tools lifts the score from 79.5% to 92.8%, highlighting the critical role of image-processing capability for high-precision GUI grounding.
Multimodal: CharXiv Reasoning
Multimodal2,323 real figures from arXiv papers across 8 scientific domains. Tests whether models can integrate visual information for multi-step reasoning — the closest existing multimodal benchmark to real scientific reading.
Mythos Preview without tools (86.1%) outperforms Opus 4.6 with tools (78.9%) by 7.2 pp — a striking generational gap.
Computer Automation: OSWorld
Computer UseAgents operate inside a real Ubuntu virtual machine, using mouse and keyboard to complete actual computer tasks: editing documents, browsing the web, managing files. Runs at 1080p resolution with a maximum of 100 actions per task.
Data Notes
· All Mythos Preview scores: adaptive thinking, max effort, default sampling settings, mean of 5 independent trials
· Competitor data sourced from each model's official System Card or benchmark leaderboard
· In Terminal-Bench 2.0, OpenAI used a proprietary harness; comparisons with other models are for reference only
· GraphWalks results cannot be reproduced via public APIs (exceeds the 1M-token limit)
· "—" indicates the model did not report a result on that benchmark