<!-- LLM_VERSION_INFO
FORMAT: text/markdown
CONTENT_TYPE: article
ORIGINAL_URL: https://www.vals.ai/benchmarks/gpqa
ALTERNATE_VERSION: benchmarks/gpqa/index.html (text/html)
EXTRACTION_DATE: 2026-04-17T00:42:42.933Z

This is the markdown version with text-only content (images converted to alt-text).
For rich formatting with images, request the HTML version at: benchmarks/gpqa/index.html
-->

# GPQA

Graduate-level Google-Proof Q&A benchmark evaluating models on questions that require deep reasoning.

## Systems (99)

| Model                                       | Accuracy     | Cost In/Out | Latency  |
|---------------------------------------------|--------------|-------------|----------|
| [Gemini 3.1 Pro Preview (02/26)](/content/models/google_gemini-3.1-pro-preview/index.html) | 95.45%± 1.05 | $2 / $12    | 65.76s   |
| [Gemini 3 Pro (11/25)](/content/models/google_gemini-3-pro-preview/index.html)         | 91.67%± 1.39 | $2 / $12    | 61.34s   |
| [GPT 5.2](/content/models/openai_gpt-5.2-2025-12-11/index.html)                       | 91.67%± 1.84 | $1.75 / $14 | 173.21s  |
| [GPT 5.4](/content/models/openai_gpt-5.4-2026-03-05/index.html)                      | 91.67%± 1.91 | $2.5 / $15   | 202.95s  |
| [Claude Opus 4.7](/content/models/anthropic_claude-opus-4-7/index.html)              | 89.90%± 1.56 | $5 / $25    | 199.85s  |
| [Muse Spark](/content/models/meta_muse_spark/index.html)                             | 89.65%± 1.55 | N/A         | 94.58s   |
| [Claude Opus 4.6 (Thinking)](/content/models/anthropic_claude-opus-4-6-thinking/index.html) | 89.65%± 1.19 | $5 / $25    | 319.02s  |
| [Grok 4.20 (Reasoning)](/content/models/grok_grok-4.20-0309-reasoning/index.html)   | 88.64%± 1.59 | $2 / $6     | 59.42s   |
| [Grok 4](/content/models/grok_grok-4-0709/index.html)                               | 88.13%± 1.63 | $3 / $15    | 115.52s  |
| [Gemini 3 Flash (12/25)](/content/models/google_gemini-3-flash-preview/index.html)   | 87.88%± 1.64 | $0.5 / $3   | 83.84s   |

## Key Takeaways

- In general, GPQA is largely saturated - progress on the benchmark has approached an asymptote and it seems any incremental improvements will be quite small.

## Results

### Input Cost Accuracy

- AI21 Labs
- Alibaba
- Anthropic
- Cohere
- DeepSeek
- Google
- Kimi
- Meta
- MiniMax
- Mistral
- Moonshot AI
- NVIDIA
- OpenAI
- xAI
- Xiaomi
- zAI

### Dataset and Context

The Graduate-Level Google-Proof Q&A (GPQA) benchmark ([paper](https://arxiv.org/abs/2311.12022)) is a public academic benchmark commonly used to measure models’ general question-answering performance. It evaluates language models on challenging, graduate-level questions across STEM fields. When created, the questions were specifically designed to be “Google-proof” - they require deep understanding and reasoning rather than fact recall or search.

As in the original paper, this benchmark uses two evaluation approaches:

- **Zero-shot chain-of-thought**: Models are asked to solve the problems with instruction to explain their reasoning steps.
- **Few-shot chain-of-thought**: In addition to the instructions, models are provided with 5 example questions and answers.

We focus on the “diamond” subset of 198 questions - these were questions in which expert validators answered correctly and no more than one out of three non-experts answered correctly. They are both challenging but have unambiguous answers.

## Additional Notes

### Methodology

We only used the diamond subset (198 questions) as the evaluation set - these were the questions that were most challenging for non-experts but had the highest agreement rate between experts. We replicated the original paper’s prompting and response parsing techniques as closely as possible, although for zero-shot CoT, we only prompt the model once, rather than asking the question across two separate queries. Here are the prompts used:

#### Zero-shot CoT Prompt:

```
What is the correct answer to this question:

Two quantum states with energies E1 and E2 have a lifetime of 10^-9 sec and 10^-8 sec, respectively. We want to clearly distinguish these two energy levels. Which one of the following options could be their energy difference so that they can be clearly resolved?

Choices:
(A) 10^-8 eV
(B) 10^-9 eV
(C) 10^-4 eV
(D) 10^-11 eV

Reason through your answer step-by-step. Then, based on your reasoning, provide the single most likely answer choice. Answer in the format "The correct answer is (insert answer here)."
```

### Few-shot CoT Prompt:

```
Here are some example questions from experts. An explanation is given before the final answer. Answer the final question yourself, giving your reasoning beforehand.
...
```

This structured content retains key information regarding GPQA, the models analyzed, the evaluation methodology, and prompts used while omitting non-semantic elements.
