CorpFin (v2)

Proprietary

Vals AI original benchmark crafted with industry experts on non-public datasets

Updated: 4/16/2026

A private benchmark evaluating understanding of long-context credit agreements

Systems Performance

Model Accuracy Cost In/Out Latency
Kimi K2.5 68.26%± 0.92 $0.6 / $3 80.78s
Qwen 3 Max Thinking 68.03%± 0.92 $1.2 / $6 121.32s
Claude Opus 4.6 (Thinking) 67.02%± 0.93 $5 / $25 20.58s
Grok 4 Fast (Reasoning) 66.90%± 0.92 $0.2 / $0.5 11.84s
Gemini 3 Flash (12/25) 66.43%± 0.93 $0.5 / $3 11.28s
Claude Opus 4.7 66.08%± 0.93 $5 / $25 17.75s
Grok 4 66.05%± 0.93 $3 / $15 29.73s
Grok 4.1 Fast (Reasoning) 65.97%± 0.93 $0.2 / $0.5 28.41s
GPT 5.2 65.89%± 0.93 $1.75 / $14 26.04s
Claude Sonnet 4.6 65.31%± 0.94 $3 / $15 18.48s

Key Takeaways

  • Kimi K2.5 leads the pack, followed closely by Qwen 3 Max Thinking.
  • Some models struggle with large context windows, particularly on the Max Fitting Context task.

Dataset and Context

In both the finance and legal industries, it is common to extract information from lengthy documents, such as credit agreements, often over 200 pages. This dataset was created with input from experts to include a series of questions about these agreements:

  • Public Validation: 20 questions from 1 document, available on request.
  • Private Validation: 340 questions from 17 documents, available for purchase.
  • Test: 858 questions from 43 documents, used for internal benchmarking.

Context Tasks

  1. Exact Pages: Provides only the necessary pages for each question.
  2. Shared Max Context: Uses a subset of pages that fit in context windows across models.
  3. Max Fitting Context: Includes the largest section of the document that fits in the model’s context window.

Types of Questions

  • Basic extraction (e.g., borrower’s legal counsel)
  • Summarization/interpretation (e.g., erroneous payment provisions)
  • Numeric reasoning (e.g., initial debt capacity)
  • References to previously defined terms
  • Market standard evaluations (e.g., unusual EBITDA terms)

Results

The performance gap between flagship and smaller models narrows when considering the cost and accuracy.

Example Question

What is the Total Net Leverage Ratio limit for unlimited RPs and investments?

Expected Answer: 3.25 to 1.00

Additional Notes

  • Context Window: Significant in influencing model performance.
  • Evaluation Methodology: Conducted using Claude 4.5 Sonnet as the judging model.