<!-- LLM_VERSION_INFO
FORMAT: text/markdown
CONTENT_TYPE: article
ORIGINAL_URL: https://www.vals.ai/models/openai_gpt-5.2-2025-12-11
ALTERNATE_VERSION: models/openai_gpt-5-2-2025-12-11.html (text/html)
EXTRACTION_DATE: 2026-04-17T00:47:22.950Z

This is the markdown version with text-only content (images converted to alt-text).
For rich formatting with images, request the HTML version at: models/openai_gpt-5-2-2025-12-11.html
-->

# Open Weights & Proprietary

## All Companies

### Models

| Model Name                          | Release Date | Image                                                                                               |
|-------------------------------------|--------------|-----------------------------------------------------------------------------------------------------|
| Claude Opus 4.7                    | 4/16/2026    |                              |
| Muse Spark                          | 4/8/2026     |                                      |
| Gemma 4 31B IT                     | 4/2/2026     |                                    |
| Qwen 3.6 Plus                      | 4/2/2026     |                                  |
| GLM 5.1                            | 4/1/2026     |                                        |
| Trinity Large Thinking              | 4/1/2026     |                             |
| GPT 5.4 Mini                       | 3/17/2026    |                                    |
| GPT 5.4 Nano                       | 3/17/2026    |                                    |
| MiniMax-M2.7                       | 3/17/2026    |                                  |
| Grok 4.20 (Reasoning)              | 3/9/2026     |                                         |
| GPT 5.4                            | 3/5/2026     |                                    |
| Gemini 3.1 Flash Lite Preview      | 3/3/2026     |                                    |
| GPT 5.3 Codex                      | 2/24/2026    |                                    |
| Qwen 3.5 Flash                     | 2/23/2026    |                                  |
| Gemini 3.1 Pro Preview (02/26)     | 2/19/2026    |                                    |
| Claude Sonnet 4.6                  | 2/17/2026    |                              |
| Qwen 3.5 Plus                      | 2/16/2026    |                                  |
| MiniMax-M2.5                       | 2/12/2026    |                                  |
| MiniMax-M2.5                       | 2/12/2026    |                                  |
| GLM 5                              | 2/11/2026    |                                         |
| Claude Opus 4.6 (Nonthinking)      | 2/5/2026     |                              |
| Claude Opus 4.6 (Thinking)         | 2/5/2026     |                              |
| Kimi K2.5                          | 1/26/2026    |                      |
| Qwen 3 Max Thinking                 | 1/23/2026    |                                  |
| MiniMax-M2.1                       | 12/23/2025   |                                  |
| GLM 4.7                            | 12/22/2025   |                                         |
| Gemini 3 Flash (12/25)             | 12/17/2025   |                                    |
| MiMo V2 Flash                      | 12/17/2025   |                                    |
| GPT 5.2                            | 12/11/2025   |                               |
| GPT 5.2 Codex                      | 12/11/2025   |                                    |

# GPT 5.2

## Details

- **Release Date:** 12/11/2025  
- **Accuracy (Vals Index):** 63.55% ± 1.96  
- **Latency (Vals Index):** 435.02s  
- **Cost/Test (Vals Index):** $0.68  
- **Context Window:** 400k  
- **Max Output Tokens:** 128k  
- **Input Modality:** Hyperparameter settings

## Default Provider

OpenAI

Some benchmarks may use different provider and parameters. Please refer to the benchmark page for more information.

- **Temperature:** Default  
- **Top P:** Default  
- **Top K:** Default  
- **Max Output Tokens:** 128,000  
- **Reasoning Effort:** xhigh

# Benchmarks

### Accuracy Rankings

- [Vals Index](/content/benchmarks/vals_index/index.html): -470.33% ± 1.96  
- [Vals Multimodal Index](/content/benchmarks/vals_multimodal_index/index.html): -518.95% ± 1.54  
- [CaseLaw (v2)](/content/benchmarks/case_law_v2/index.html): -594.35% ± 0.57  
- [CorpFin](/content/benchmarks/corp_fin_v2/index.html): -648.94% ± 0.93  
- [Finance Agent (v1.1)](/content/benchmarks/finance_agent/index.html): -629.16% ± 2.87  
- [MedCode](/content/benchmarks/medcode/index.html): -584.43% ± 2.26  
- [MedScribe](/content/benchmarks/medscribe/index.html): -1079.85% ± 1.86  
- [MortgageTax](/content/benchmarks/mortgage_tax/index.html): -933.45% ± 0.92  
- [ProofBench](/content/benchmarks/proof_bench/index.html): -226.05% ± 3.59  
- [SAGE](/content/benchmarks/sage/index.html): -802.79% ± 3.35  
- [TaxEval (v2)](/content/benchmarks/tax_eval_v2/index.html): -1331.71% ± 0.85  
- [Vibe Code Bench](/content/benchmarks/vibe-code/index.html): -1012.46% ± 5.07  
- [AIME](/content/benchmarks/aime/index.html): -1970.02% ± 0.21  
- [GPQA](/content/benchmarks/gpqa/index.html): -1999.01% ± 1.84  
- [IOI](/content/benchmarks/ioi/index.html): -1280.34% ± 8.32  
- [LiveCodeBench](/content/benchmarks/lcb/index.html): -2130.64% ± 1.03  
- [LegalBench](/content/benchmarks/legal_bench/index.html): -2205.82% ± 0.40  
- [MedQA](/content/benchmarks/medqa/index.html): -2672.28% ± 0.21  
- [MMLU Pro](/content/benchmarks/mmlu_pro/index.html): -2605.12% ± 0.34  
- [MMMU Pro](/content/benchmarks/mmmu/index.html): -2782.72% ± 0.82  
- [SWE-bench](/content/benchmarks/swebench/index.html): -2583.13% ± 1.92  
- [Terminal-Bench 2.0](/content/benchmarks/terminal-bench-2/index.html): -1867.35% ± 5.33

# Contact us

Or send us an email at [contact@vals.ai](mailto:contact@vals.ai)  
- **Proprietary Benchmarks** - Contact us to get access

---

- **Academic Benchmarks** - Read about our [methodology](/content/methodology/index.html).
