<!-- LLM_VERSION_INFO
FORMAT: text/markdown
CONTENT_TYPE: article
ORIGINAL_URL: https://www.vals.ai/models/google_gemini-2.5-pro
ALTERNATE_VERSION: models/google_gemini-2-5-pro.html (text/html)
EXTRACTION_DATE: 2026-04-17T00:42:59.819Z

This is the markdown version with text-only content (images converted to alt-text).
For rich formatting with images, request the HTML version at: models/google_gemini-2-5-pro.html
-->

# Open Weights & Proprietary

## All Companies

### Models

#### Release date

| Model Name | Release Date | Image |
| --- | --- | --- |
| Claude Opus 4.7 | 4/16/2026 |  |
| Muse Spark | 4/8/2026 |  |
| Gemma 4 31B IT | 4/2/2026 |  |
| Qwen 3.6 Plus | 4/2/2026 |  |
| GLM 5.1 | 4/1/2026 |  |
| Trinity Large Thinking | 4/1/2026 |  |
| GPT 5.4 Mini | 3/17/2026 |  |
| GPT 5.4 Nano | 3/17/2026 |  |
| MiniMax-M2.7 | 3/17/2026 |  |
| Grok 4.20 (Reasoning) | 3/9/2026 |  |
| GPT 5.4 | 3/5/2026 |  |
| Gemini 3.1 Flash Lite Preview | 3/3/2026 |  |
| GPT 5.3 Codex | 2/24/2026 |  |
| Qwen 3.5 Flash | 2/23/2026 |  |
| Gemini 3.1 Pro Preview (02/26) | 2/19/2026 |  |
| Claude Sonnet 4.6 | 2/17/2026 |  |
| Qwen 3.5 Plus | 2/16/2026 |  |
| MiniMax-M2.5 | 2/12/2026 |  |
| GLM 5 | 2/11/2026 |  |
| Claude Opus 4.6 (Nonthinking) | 2/5/2026 |  |
| Claude Opus 4.6 (Thinking) | 2/5/2026 |  |
| Kimi K2.5 | 1/26/2026 |  |
| Qwen 3 Max Thinking | 1/23/2026 |  |
| MiniMax-M2.1 | 12/23/2025 |  |
| GLM 4.7 | 12/22/2025 |  |
| Gemini 3 Flash (12/25) | 12/17/2025 |  |
| MiMo V2 Flash | 12/17/2025 |  |
| GPT 5.2 | 12/11/2025 |  |
| GPT 5.2 Codex | 12/11/2025 |  |

### [Gemini 2.5 Pro](https://ai.google.dev/gemini-api/docs/models)

- **Release Date**: 7/17/2025  
- **Default Provider**: Google

#### Accuracy (Vals Index)
- **Value**: 48.82%
- **Margin of Error**: ± 1.99

#### Latency (Vals Index)
- **Value**: 588.73s

#### Cost/Test (Vals Index)
- **Amount**: $0.31

#### Context Window
- **Value**: 1M

#### Max Output Tokens
- **Amount**: 66k

#### Input Modality
- **Hyperparameter settings**
  - Temperature: 1
  - Top P: Default
  - Top K: Default
  - Max Output Tokens: 65,536

### Benchmarks

- **Accuracy Rankings**
  - [Vals Index](/content/benchmarks/vals_index/index.html)
    - **Value**: -119.10% ± 1.99
    - **Rank**: 55/40
  - [Vals Multimodal Index](/content/benchmarks/vals_multimodal_index/index.html)
    - **Value**: -143.48% ± 1.56
    - **Rank**: 42/28
  - [CaseLaw (v2)](/content/benchmarks/case_law_v2/index.html)
    - **Value**: -209.47% ± 0.60
    - **Rank**: 158/47
  - [CorpFin](/content/benchmarks/corp_fin_v2/index.html)
    - **Value**: -230.53% ± 0.96
    - **Rank**: 332/97
  - [Finance Agent (v1.1)](/content/benchmarks/finance_agent/index.html)
    - **Value**: -180.69% ± 2.78
    - **Rank**: 97/45
  - [MedCode](/content/benchmarks/medcode/index.html)
    - **Value**: -249.89% ± 2.11
    - **Rank**: 268/51
  - [MedScribe](/content/benchmarks/medscribe/index.html)
    - **Value**: -410.02% ± 1.91
    - **Rank**: 146/51
  - [MortgageTax](/content/benchmarks/mortgage_tax/index.html)
    - **Value**: -431.19% ± 0.91
    - **Rank**: 476/69
  - [SAGE](/content/benchmarks/sage/index.html)
    - **Value**: -292.81% ± 3.41
    - **Rank**: 231/49
  - [IOI](/content/benchmarks/ioi/index.html)
    - **Value**: -132.45% ± 6.17
    - **Rank**: 314/50
  - [SWE-bench](/content/benchmarks/swebench/index.html)
    - **Value**: -466.50% ± 2.23
    - **Rank**: 84/41
  - [Terminal-Bench 2.0](/content/benchmarks/terminal-bench-2/index.html)
    - **Value**: -286.65% ± 4.90
    - **Rank**: 241/52

### Contact us
- **Email**: contact@vals.ai  
- **Message**: Proprietary Benchmarks (contact us for access)  
- **Methodology**: Read about our [methodology](/content/methodology/index.html).
