<!-- LLM_VERSION_INFO
FORMAT: text/markdown
CONTENT_TYPE: article
ORIGINAL_URL: https://www.vals.ai/models/openai_gpt-5.1-2025-11-13
ALTERNATE_VERSION: models/openai_gpt-5-1-2025-11-13.html (text/html)
EXTRACTION_DATE: 2026-04-17T00:43:54.205Z

This is the markdown version with text-only content (images converted to alt-text).
For rich formatting with images, request the HTML version at: models/openai_gpt-5-1-2025-11-13.html
-->

## Open Weights & Proprietary

### All Companies

| Release Date | Model Name | Provider |  
|---------------|------------|----------|  
| 4/16/2026     | Claude Opus 4.7        |       |  
| 4/8/2026      | Muse Spark             |             |  
| 4/2/2026      | Gemma 4 31B IT        |        |  
| 4/2/2026      | Qwen 3.6 Plus         |      |  
| 4/1/2026      | GLM 5.1               |              |  
| 4/1/2026      | Trinity Large Thinking  |  |  
| 3/17/2026     | GPT 5.4 Mini          |        |  
| 3/17/2026     | GPT 5.4 Nano          |        |  
| 3/17/2026     | MiniMax-M2.7          |      |  
| 3/9/2026      | Grok 4.20 (Reasoning) |              |  
| 3/5/2026      | GPT 5.4               |        |  
| 3/3/2026      | Gemini 3.1 Flash Lite Preview |   |  
| 2/24/2026     | GPT 5.3 Codex         |        |  
| 2/23/2026     | Qwen 3.5 Flash        |      |  
| 2/19/2026     | Gemini 3.1 Pro Preview (02/26) |   |  
| 2/17/2026     | Claude Sonnet 4.6     |  |  
| 2/16/2026     | Qwen 3.5 Plus         |      |  
| 2/12/2026     | MiniMax-M2.5          |      |  
| 2/12/2026     | MiniMax-M2.5          |      |  
| 2/11/2026     | GLM 5                 |              |  
| 2/5/2026      | Claude Opus 4.6 (Nonthinking) |  |  
| 2/5/2026      | Claude Opus 4.6 (Thinking) |  |  
| 1/26/2026     | Kimi K2.5             |  |  
| 1/23/2026     | Qwen 3 Max Thinking    |      |  
| 12/23/2025    | MiniMax-M2.1          |      |  
| 12/22/2025    | GLM 4.7               |              |  
| 12/17/2025    | Gemini 3 Flash (12/25)|        |  
| 12/17/2025    | MiMo V2 Flash         |        |  
| 12/11/2025    | GPT 5.2               |        |  
| 12/11/2025    | GPT 5.2 Codex         |        |

### GPT 5.1 Details

**Release Date:** 11/13/2025  
**Provider:**

- **Accuracy (Vals Index):** 60.38% ± 2.00  
- **Latency (Vals Index):** 411.84s  
- **Cost/Test (Vals Index):** $0.36  
- **Context Window:** 400k  
- **Max Output Tokens:** 128k

### Hyperparameter Settings

- **Temperature:** Default  
- **Top P:** Default  
- **Top K:** Default  
- **Max Output Tokens:** 128,000  
- **Reasoning Effort:** High

### Benchmarks  
  
- **Accuracy Rankings:** -73.04% ± 2.00  
- **Vals Multimodal Index:** -89.68% ± 1.57  
- **CaseLaw (v2):** -138.36% ± 0.75  
- **CorpFin:** -145.62% ± 0.95  
- **Finance Agent (v1.1):** -150.06% ± 2.80  
- **MedCode:** -167.74% ± 2.15  
- **MedScribe:** -324.77% ± 1.94  
- **MortgageTax:** -259.77% ± 0.93  
- **SAGE:** -208.31% ± 3.16  
- **TaxEval (v2):** -407.63% ± 0.85  
- **AIME:** -570.87% ± 0.62  
- **GPQA:** -591.89% ± 1.71  
- **IOI:** -163.35% ± 7.34  
- **LiveCodeBench:** -727.29% ± 0.98  
- **LegalBench:** -801.22% ± 0.40  
- **MedQA:** -981.26% ± 0.17  
- **MMLU Pro:** -962.55% ± 0.34  
- **MMMU Pro:** -1011.55% ± 0.90  
- **SWE-bench:** -923.74% ± 2.06  
- **Terminal-Bench 2.0:** -645.52% ± 5.30

### Contact Information

For proprietary benchmarks or general inquiries, please email [contact@vals.ai](mailto:contact@vals.ai).  
Read about our [methodology](/content/methodology/index.html).
