CorpFin (v2)
Proprietary
Vals AI original benchmark crafted with industry experts on non-public datasets
Updated: 4/16/2026
A private benchmark evaluating understanding of long-context credit agreements
Systems Performance
| Model | Accuracy | Cost In/Out | Latency |
|---|---|---|---|
| Kimi K2.5 | 68.26%± 0.92 | $0.6 / $3 | 80.78s |
| Qwen 3 Max Thinking | 68.03%± 0.92 | $1.2 / $6 | 121.32s |
| Claude Opus 4.6 (Thinking) | 67.02%± 0.93 | $5 / $25 | 20.58s |
| Grok 4 Fast (Reasoning) | 66.90%± 0.92 | $0.2 / $0.5 | 11.84s |
| Gemini 3 Flash (12/25) | 66.43%± 0.93 | $0.5 / $3 | 11.28s |
| Claude Opus 4.7 | 66.08%± 0.93 | $5 / $25 | 17.75s |
| Grok 4 | 66.05%± 0.93 | $3 / $15 | 29.73s |
| Grok 4.1 Fast (Reasoning) | 65.97%± 0.93 | $0.2 / $0.5 | 28.41s |
| GPT 5.2 | 65.89%± 0.93 | $1.75 / $14 | 26.04s |
| Claude Sonnet 4.6 | 65.31%± 0.94 | $3 / $15 | 18.48s |
Key Takeaways
- Kimi K2.5 leads the pack, followed closely by Qwen 3 Max Thinking.
- Some models struggle with large context windows, particularly on the Max Fitting Context task.
Dataset and Context
In both the finance and legal industries, it is common to extract information from lengthy documents, such as credit agreements, often over 200 pages. This dataset was created with input from experts to include a series of questions about these agreements:
- Public Validation: 20 questions from 1 document, available on request.
- Private Validation: 340 questions from 17 documents, available for purchase.
- Test: 858 questions from 43 documents, used for internal benchmarking.
Context Tasks
- Exact Pages: Provides only the necessary pages for each question.
- Shared Max Context: Uses a subset of pages that fit in context windows across models.
- Max Fitting Context: Includes the largest section of the document that fits in the model’s context window.
Types of Questions
- Basic extraction (e.g., borrower’s legal counsel)
- Summarization/interpretation (e.g., erroneous payment provisions)
- Numeric reasoning (e.g., initial debt capacity)
- References to previously defined terms
- Market standard evaluations (e.g., unusual EBITDA terms)
Results
The performance gap between flagship and smaller models narrows when considering the cost and accuracy.
Example Question
What is the Total Net Leverage Ratio limit for unlimited RPs and investments?
Expected Answer: 3.25 to 1.00
Additional Notes
- Context Window: Significant in influencing model performance.
- Evaluation Methodology: Conducted using Claude 4.5 Sonnet as the judging model.