<!-- LLM_VERSION_INFO
FORMAT: text/markdown
CONTENT_TYPE: article
ORIGINAL_URL: https://www.vals.ai/models/openai_gpt-5-2025-08-07
ALTERNATE_VERSION: models/openai_gpt-5-2025-08-07/index.html (text/html)
EXTRACTION_DATE: 2026-04-17T00:47:21.335Z

This is the markdown version with text-only content (images converted to alt-text).
For rich formatting with images, request the HTML version at: models/openai_gpt-5-2025-08-07/index.html
-->

# Open Weights & Proprietary

## Release Dates and Models

| Release Date | Model Name | Image |
|--------------|------------|-------|
| 4/16/2026    | Claude Opus 4.7        |  |
| 4/8/2026     | Muse Spark             |  |
| 4/2/2026     | Gemma 4 31B IT        |  |
| 4/2/2026     | Qwen 3.6 Plus         |  |
| 4/1/2026     | GLM 5.1               |  |
| 4/1/2026     | Trinity Large Thinking  |  |
| 3/17/2026    | GPT 5.4 Mini          |  |
| 3/17/2026    | GPT 5.4 Nano          |  |
| 3/17/2026    | MiniMax-M2.7          |  |
| 3/9/2026     | Grok 4.20 (Reasoning) |  |
| 3/5/2026     | GPT 5.4               |  |
| 3/3/2026     | Gemini 3.1 Flash Lite Preview |  |
| 2/24/2026    | GPT 5.3 Codex         |  |
| 2/23/2026    | Qwen 3.5 Flash        |  |
| 2/19/2026    | Gemini 3.1 Pro Preview (02/26) |  |
| 2/17/2026    | Claude Sonnet 4.6     |  |
| 2/16/2026    | Qwen 3.5 Plus         |  |
| 2/12/2026    | MiniMax-M2.5          |  |
| 2/12/2026    | MiniMax-M2.5          |  |
| 2/11/2026    | GLM 5                 |  |
| 2/5/2026     | Claude Opus 4.6 (Nonthinking) |  |
| 2/5/2026     | Claude Opus 4.6 (Thinking) |  |
| 1/26/2026    | Kimi K2.5             |  |
| 1/23/2026    | Qwen 3 Max Thinking    |  |
| 12/23/2025   | MiniMax-M2.1          |  |
| 12/22/2025   | GLM 4.7               |  |
| 12/17/2025   | Gemini 3 Flash (12/25) |  |
| 12/17/2025   | MiMo V2 Flash         |  |
| 12/11/2025    | GPT 5.2               |  |
| 12/11/2025    | GPT 5.2 Codex         |  |

## GPT 5
- **Release Date:** 8/7/2025  
- **Default Provider:** OpenAI  
**Accuracy (Vals Index):** 56.10% ± 2.00  
**Latency (Vals Index):** 513.66s  
**Cost/Test (Vals Index):** $0.29  
**Context Window:** 400k  
**Max Output Tokens:** 128k  
**Input Modality:** Hyperparameter settings

### Benchmarks

#### Accuracy Rankings
- [Vals Index](/content/benchmarks/vals_index/index.html) -72.78% ± 2.00
- [Vals Multimodal Index](/content/benchmarks/vals_multimodal_index/index.html) -91.63% ± 1.56
- [CaseLaw (v2)](/content/benchmarks/case_law_v2/index.html) -132.69% ± 0.65
- [CorpFin](/content/benchmarks/corp_fin_v2/index.html) -146.75% ± 0.96
- [Finance Agent (v1.1)](/content/benchmarks/finance_agent/index.html) -148.41% ± 2.88
- [MedCode](/content/benchmarks/medcode/index.html) -164.96% ± 2.10
- [MedScribe](/content/benchmarks/medscribe/index.html) -321.28% ± 1.94
- [MortgageTax](/content/benchmarks/mortgage_tax/index.html) -287.81% ± 0.88
- [ProofBench](/content/benchmarks/proof_bench/index.html) -89.91% ± 3.86
- [SAGE](/content/benchmarks/sage/index.html) -246.15% ± 3.35
- [TaxEval (v2)](/content/benchmarks/tax_eval_v2/index.html) -463.81% ± 0.87
- [AIME](/content/benchmarks/aime/index.html) -658.22% ± 1.31
- [GPQA](/content/benchmarks/gpqa/index.html) -670.02% ± 1.76
- [IOI](/content/benchmarks/ioi/index.html) -173.07% ± 4.58
- [LiveCodeBench](/content/benchmarks/lcb/index.html) -820.93% ± 0.97
- [LegalBench](/content/benchmarks/legal_bench/index.html) -899.20% ± 0.38
- [MedQA](/content/benchmarks/medqa/index.html) -1101.20% ± 0.17
- [MMLU Pro](/content/benchmarks/mmlu_pro/index.html) -1078.84% ± 0.34
- [MMMU Pro](/content/benchmarks/mmmu/index.html) -1104.98% ± 0.93
- [SWE-bench](/content/benchmarks/swebench/index.html) -1014.47% ± 2.07
- [Terminal-Bench 2.0](/content/benchmarks/terminal-bench-2/index.html) -590.04% ± 5.15

## Contact Information
Or send us an email at [contact@vals.ai](mailto:contact@vals.ai).

## Proprietary Benchmarks
[Contact us](/content/models/openai_gpt-5-2025-08-07#contact-form/index.html) to get access.

[Read about our methodology.](/content/methodology/index.html)
