<!-- LLM_VERSION_INFO
FORMAT: text/markdown
CONTENT_TYPE: article
ORIGINAL_URL: https://www.vals.ai/models/openai_gpt-5.4-2026-03-05
ALTERNATE_VERSION: models/openai_gpt-5-4-2026-03-05.html (text/html)
EXTRACTION_DATE: 2026-04-17T00:42:48.412Z

This is the markdown version with text-only content (images converted to alt-text).
For rich formatting with images, request the HTML version at: models/openai_gpt-5-4-2026-03-05.html
-->

## Open Weights & Proprietary

### Models

| Model Name                             | Release Date | Image                                                                                           |
|----------------------------------------|--------------|-------------------------------------------------------------------------------------------------|
| Claude Opus 4.7                       | 4/16/2026    |                          |
| Muse Spark                             | 4/8/2026     |                                    |
| Gemma 4 31B IT                        | 4/2/2026     |                                |
| Qwen 3.6 Plus                         | 4/2/2026     |                              |
| GLM 5.1                               | 4/1/2026     |                                      |
| Trinity Large Thinking                 | 4/1/2026     |                         |
| GPT 5.4 Mini                          | 3/17/2026    |                                |
| GPT 5.4 Nano                          | 3/17/2026    |                                |
| MiniMax-M2.7                          | 3/17/2026    |                              |
| Grok 4.20 (Reasoning)                 | 3/9/2026     |                                      |
| GPT 5.4                               | 3/5/2026     |                           |
| Gemini 3.1 Flash Lite Preview         | 2/24/2026    |                                |
| GPT 5.3 Codex                         | 2/23/2026    |                                |
| Qwen 3.5 Flash                        | 2/19/2026    |                              |
| Gemini 3.1 Pro Preview (02/26)       | 2/17/2026    |                                |
| Claude Sonnet 4.6                     | 2/16/2026    |                          |
| Qwen 3.5 Plus                         | 2/12/2026    |                              |
| MiniMax-M2.5                          | 2/12/2026    |                              |
| MiniMax-M2.5                          | 2/11/2026    |                              |
| GLM 5                                 | 2/5/2026     |                                      |
| Claude Opus 4.6 (Nonthinking)        | 2/5/2026     |                          |
| Claude Opus 4.6 (Thinking)           | 1/26/2026    |                          |
| Kimi K2.5                             | 1/23/2026    |                  |
| Qwen 3 Max Thinking                   | 12/23/2025   |                              |
| MiniMax-M2.1                          | 12/22/2025   |                              |
| GLM 4.7                               | 12/17/2025   |                                      |
| Gemini 3 Flash (12/25)               | 12/17/2025   |                                |
| MiMo V2 Flash                         | 12/11/2025   |                                |
| GPT 5.2                               | 12/11/2025   |                                |
| GPT 5.2 Codex                         | 12/11/2025   |                                |

### GPT 5.4

- **Release Date:** 3/5/2026
- **Accuracy (Vals Index):** 64.77% ± 1.95
- **Latency (Vals Index):** 460.70s
- **Cost/Test (Vals Index):** $0.67
- **Context Window:** 1M
- **Max Output Tokens:** 128k
- **Input Modality:** Hyperparameter settings

#### Default Provider:
OpenAI

### Some benchmarks may use different provider and parameters. Please refer to the benchmark page for more information.

- **Temperature:** Default
- **Top P:** Default
- **Top K:** Default
- **Max Output Tokens:** 128,000
- **Reasoning Effort:** xhigh

## Benchmarks

### Accuracy Rankings

- [Vals Index](/content/benchmarks/vals_index/index.html)  -139.19% ± 1.95 113/40
- [Vals Multimodal Index](/content/benchmarks/vals_multimodal_index/index.html) -162.37% ± 1.52 86/28
- [CaseLaw (v2)](/content/benchmarks/case_law_v2/index.html) -190.76% ± 2.24 146/47
- [CorpFin](/content/benchmarks/corp_fin_v2/index.html) -227.27% ± 0.94 393/97
- [Finance Agent (v1.1)](/content/benchmarks/finance_agent/index.html) -229.30% ± 2.85 189/45
- [MedCode](/content/benchmarks/medcode/index.html) -189.13% ± 2.15 188/51
- [MedScribe](/content/benchmarks/medscribe/index.html) -402.70% ± 3.32 191/51
- [MortgageTax](/content/benchmarks/mortgage_tax/index.html) -399.45% ± 0.91 426/69
- [ProofBench](/content/benchmarks/proof_bench/index.html) -366.46% ± 4.99 175/24
- [SAGE](/content/benchmarks/sage/index.html) -315.66% ± 3.12 275/49
- [TaxEval (v2)](/content/benchmarks/tax_eval_v2/index.html) -597.62% ± 0.87 767/104
- [Vibe Code Bench](/content/benchmarks/vibe-code/index.html) -601.46% ± 4.84 240/26
- [AIME](/content/benchmarks/aime/index.html) -948.54% ± 0.53 989/96
- [GPQA](/content/benchmarks/gpqa/index.html) -986.64% ± 1.91 1121/99
- [IOI](/content/benchmarks/ioi/index.html) -797.10% ± 9.87 626/50
- [LiveCodeBench](/content/benchmarks/lcb/index.html) -1077.38% ± 1.04 1217/103
- [LegalBench](/content/benchmarks/legal_bench/index.html) -1196.79% ± 0.41 1674/116
- [MedQA](/content/benchmarks/medqa/index.html) -1448.74% ± 0.18 1452/95
- [MMLU Pro](/content/benchmarks/mmlu_pro/index.html) -1426.02% ± 0.42 1548/97
- [MMMU Pro](/content/benchmarks/mmmu/index.html) -1538.71% ± 0.79 1156/66
- [SWE-bench](/content/benchmarks/swebench/index.html) -1480.55% ± 1.85 761/41
- [Terminal-Bench 2.0](/content/benchmarks/terminal-bench-2/index.html) -1188.49% ± 5.25 947/52

### Contact us

- Or send us an email at [contact@vals.ai](mailto:contact@vals.ai)
- Proprietary Benchmarks ([contact us](/content/models/openai_gpt-5.4-2026-03-05#contact-form/index.html) to get access)
- Read about our [methodology](/content/methodology/index.html).
