Open Weights & Proprietary
All Companies
Release date Models
4/16/2026
Claude Opus 4.7
4/8/2026
Muse Spark
4/2/2026
Gemma 4 31B IT
4/2/2026
Qwen 3.6 Plus
4/1/2026
GLM 5.1
4/1/2026
Trinity Large Thinking
3/17/2026
GPT 5.4 Mini
3/17/2026
GPT 5.4 Nano
3/17/2026
MiniMax-M2.7
3/9/2026
Grok 4.20 (Reasoning)
3/5/2026
GPT 5.4
3/3/2026
Gemini 3.1 Flash Lite Preview
2/24/2026
GPT 5.3 Codex
2/23/2026
Qwen 3.5 Flash
2/19/2026
Gemini 3.1 Pro Preview (02/26)
2/17/2026
Claude Sonnet 4.6
2/16/2026
Qwen 3.5 Plus
2/12/2026
MiniMax-M2.5
2/12/2026
MiniMax-M2.5
2/11/2026
GLM 5
2/5/2026
Claude Opus 4.6 (Nonthinking)
2/5/2026
Claude Opus 4.6 (Thinking)
1/26/2026
Kimi K2.5
1/23/2026
Qwen 3 Max Thinking
12/23/2025
MiniMax-M2.1
12/22/2025
GLM 4.7
12/17/2025
Gemini 3 Flash (12/25)
12/17/2025
MiMo V2 Flash
12/11/2025
GPT 5.2
12/11/2025
GPT 5.2 Codex
Claude Opus 4.6 (Thinking)
Compare Models
Claude Opus 4.6 (Thinking)
Release Date: 2/5/2026
Accuracy (Vals Index)
65.88%
± 1.94
Latency (Vals Index)
334.54s
Cost/Test (Vals Index)
$0.89
Context Window
200k
Max Output Tokens
128k
Input Modality
Hyperparameter settings
- Default Provider: Anthropic
Some benchmarks may use different provider and parameters. Please refer to the benchmark page for more information.
| Temperature | 1 |
|---|---|
| Top P | Default |
| Top K | Default |
| Max Output Tokens | 128,000 |
| Compute Effort | max |
Benchmarks
Accuracy
Rankings
Vals Index](/content/benchmarks/vals_index/index.html) -72.44%
± 1.94
81/ 40Vals Multimodal Index](/content/benchmarks/vals_multimodal_index/index.html) -91.32%
± 1.53
62/ 28CaseLaw (v2) -108.87%
± 0.37
98/ 47CorpFin -143.38%
± 0.93
298/ 97Finance Agent (v1.1) -153.66%
± 2.78
150/ 45MedCode -148.08%
± 2.09
172/ 51MedScribe -302.01%
± 1.94
219/ 51MortgageTax -276.66%
± 0.91
319/ 69ProofBench -230.45%
± 5.03
121/ 24SAGE -269.93%
± 3.34
279/ 49TaxEval (v2) -447.11%
± 0.83
699/ 104Vibe Code Bench -351.96%
± 4.68
158/ 26AIME -700.67%
± 0.64
733/ 96GPQA -727.91%
± 1.19
846/ 99LiveCodeBench -758.96%
± 1.02
901/ 103LegalBench -840.74%
± 0.37
1191/ 116MedQA -1030.59%
± 0.19
992/ 95MMLU Pro -1051.55%
± 0.45
1195/ 97MMMU Pro -1078.09%
± 0.88
786/ 66SWE-bench -1092.00%
± 1.85
558/ 41Terminal-Bench 2.0 -884.16%
± 5.25
733/ 52
Contact us
- Or send us an email at contact@vals.ai
Proprietary Benchmarks
Academic Benchmarks
Read about our methodology.