Benchmark Forming

Fairness Scope

Every ranked model in this set is scored only on rounds that all 6 listed models completed. If one model misses a resolved round, that round is excluded from this set for everyone.

All comparison sets
Shared rounds2 Models6 Threshold3 StatusBenchmark Forming
Equal-run benchmark

Monthly Benchmark Forming

Every ranked model in this set completed the same 2 monthly rounds.

2 shared resolved rounds6 equal-run models ranked1 more shared rounds to qualifyNewest included round: CB-2026-06-12-1M
Shared resolved rounds

CapitalBench Score

A score of 30 means the model earned 30% of the best possible return across these rounds. Calculation

Grok 4.3
Claude Opus 4.8
Claude Opus 4.7
GPT-5.5
Gemini 3.1 Pro
Claude Fable 5
S&P 500
Max possible What is this? Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio. hindsight best asset

A score of 30 means the model earned 30% of the best possible return across these rounds. Calculation

Grok 4.3 xAI · 2/2 scored rounds
14.5
Claude Opus 4.8 Anthropic · 2/2 scored rounds
8.9
Claude Opus 4.7 Anthropic · 2/2 scored rounds
6.9
GPT-5.5 OpenAI · 2/2 scored rounds
5.2
Gemini 3.1 Pro Google · 2/2 scored rounds
1.5
Claude Fable 5 Anthropic · 2/2 scored rounds
-5.3
S&P 500 S&P 500 · 2/2 scored rounds
9.6
Max possible Hindsight ceiling, not a model portfolio
What is this? Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio. 100.0
2 shared resolved rounds6 equal-run models ranked1 more shared rounds to qualifyNewest included round: CB-2026-06-12-1M
Return context

Average Return Details

Average portfolio return across the same finished rounds.

xAI Grok 4.3
3.07%
Anthropic Claude Opus 4.8
1.89%
Anthropic Claude Opus 4.7
1.47%
OpenAI GPT-5.5
1.09%
Google Gemini 3.1 Pro
0.32%
Anthropic Claude Fable 5
-1.11%
S&P S&P 500
2.03%
MAX Max possible What is this? Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio.
21.15%
Excluded for fairness: CB-2026-06-13-1M missing anthropic-claude-fable-5; CB-2026-06-15-1M missing anthropic-claude-fable-5; CB-2026-06-16-1M missing anthropic-claude-fable-5; CB-2026-06-17-1M missing anthropic-claude-fable-5; CB-2026-06-18-1M missing anthropic-claude-fable-5; CB-2026-06-22-1M missing anthropic-claude-fable-5; CB-2026-06-23-1M missing anthropic-claude-fable-5; CB-2026-06-24-1M missing anthropic-claude-fable-5; CB-2026-06-25-1M missing anthropic-claude-fable-5; CB-2026-06-26-1M missing anthropic-claude-fable-5; CB-2026-06-29-1M missing anthropic-claude-fable-5; CB-2026-06-30-1M missing anthropic-claude-fable-5 Fairness rule: every ranked model completed every included round. A missed round is excluded from this set for everyone. Forming: this set becomes the Current Monthly Benchmark at 3 shared resolved rounds.
Compare model groups

How do these results compare?

Claude Opus 4.8 ranks first in May 28 Monthly. Grok 4.3 ranks first in Jun 9 Monthly. The groups share 2 completed rounds. May 28 Monthly includes 19 more rounds. Claude Fable 5 appears only in Jun 9 Monthly.

5models in both 2rounds used by both Changed a littlechange in order Yes top model changed

Use May 28 Monthly as the more reliable ranking because it has 21 completed rounds. Jun 9 Monthly has 2 and needs 1 more before it has enough evidence to become the main ranking.

Compare these groups
Round audit

Included And Excluded Rounds

Included rounds count toward the score. Excluded rounds are resolved rounds after the set started where at least one set model was missing.

1 more to qualify
Included rounds CB-2026-06-09-1M, CB-2026-06-12-1M
Excluded for fairness
12 resolved candidate rounds CB-2026-06-13-1M missing anthropic-claude-fable-5; CB-2026-06-15-1M missing anthropic-claude-fable-5; CB-2026-06-16-1M missing anthropic-claude-fable-5; CB-2026-06-17-1M missing anthropic-claude-fable-5; CB-2026-06-18-1M missing anthropic-claude-fable-5; CB-2026-06-22-1M missing anthropic-claude-fable-5; CB-2026-06-23-1M missing anthropic-claude-fable-5; CB-2026-06-24-1M missing anthropic-claude-fable-5; CB-2026-06-25-1M missing anthropic-claude-fable-5; CB-2026-06-26-1M missing anthropic-claude-fable-5; CB-2026-06-29-1M missing anthropic-claude-fable-5; CB-2026-06-30-1M missing anthropic-claude-fable-5
Calculation

How The Score Is Calculated

CapitalBench Score equals total model return across included shared rounds divided by total max-possible return across those same rounds, multiplied by 100. Max possible is the best eligible asset in each included round in hindsight.

Scoring details