Qualified Comparison Set

Fairness Scope

Every ranked model in this set is scored only on rounds that all 7 listed models completed. If one model misses a resolved round, that round is excluded from this set for everyone.

Includes 5 earlier shared rounds completed by all 7 models before this roster began, plus 10 completed since the roster changed.

All comparison sets
Shared rounds15 Models7 Threshold3 StatusQualified Comparison Set
Equal-run benchmark

Monthly Qualified Comparison Set

Every ranked model in this set completed the same 15 monthly rounds.

15 shared resolved rounds7 equal-run models rankedQualified at 3+ shared roundsNewest included round: CB-2026-08-04-1M
Shared resolved rounds

CapitalBench Score

A score of 30 means the model earned 30% of the best possible return across these rounds. Calculation

Grok 4.5
Gemini 3.1 Pro
GPT-5.5
Claude Opus 4.8
Claude Fable 5
Grok 4.3
GPT-5.6 Sol
S&P 500
Max possible What is this? Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio. hindsight best asset

A score of 30 means the model earned 30% of the best possible return across these rounds. Calculation

Grok 4.5 xAI · 15/15 scored rounds
17.1
Gemini 3.1 Pro Google · 15/15 scored rounds
16.7
GPT-5.5 OpenAI · 15/15 scored rounds
14.4
Claude Opus 4.8 Anthropic · 15/15 scored rounds
14.0
Claude Fable 5 Anthropic · 15/15 scored rounds
13.7
Grok 4.3 xAI · 15/15 scored rounds
12.9
GPT-5.6 Sol OpenAI · 15/15 scored rounds
11.9
S&P 500 S&P 500 · 15/15 scored rounds
13.2
Max possible Hindsight ceiling, not a model portfolio
What is this? Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio. 100.0
15 shared resolved rounds7 equal-run models rankedQualified at 3+ shared roundsNewest included round: CB-2026-08-04-1M
Return context

Average Return Details

Average portfolio return across the same finished rounds.

xAI Grok 4.5
4.15%
Google Gemini 3.1 Pro
4.06%
OpenAI GPT-5.5
3.49%
Anthropic Claude Opus 4.8
3.39%
Anthropic Claude Fable 5
3.33%
xAI Grok 4.3
3.14%
OpenAI GPT-5.6 Sol
2.89%
S&P S&P 500
3.20%
MAX Max possible What is this? Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio.
24.28%
Fairness rule: every ranked model completed every included round. A missed round is excluded from this set for everyone.
Risk and return

Who earned more return for the risk they took?

Average realized return and frozen portfolio risk across the same 15 shared monthly rounds.

Risk and return for models in this comparison setAverage frozen portfolio allocation risk is plotted horizontally and average realized return is plotted vertically across 15 shared rounds.-2.00%0.00%2.00%4.00%6.00%5060708090CapitalBench allocation riskAverage returnS&P 500Beat S&P with lower allocation riskGrok 4.5Gemini 3.1GPT-5.5Opus 4.8Fable 5Grok 4.3GPT-5.6 Sol
Risk and return for models in this comparison setAverage frozen portfolio allocation risk is plotted horizontally and average realized return is plotted vertically across 15 shared rounds.-2.00%0.00%2.00%4.00%6.00%5060708090CapitalBench allocation riskAverage returnS&P 500Grok 4.5Gemini 3.1GPT-5.5Opus 4.8Fable 5Grok 4.3GPT-5.6 Sol
Grok 4.5xAI
+4.15%average return69.6/100Risk-seeking+0.94 ppversus S&P 50059.5-92.8risk range

Return leaderGrok 4.5 led the models at +4.15% average return with a 69.6/100 risk score.

Benchmark testGemini 3.1 Pro and Claude Opus 4.8 beat the S&P 500 while taking no more allocation risk.

Compare model groups

How do these results compare?

Grok 4.5 ranks first in both groups. The groups share 7 completed rounds. Jul 21 Monthly includes 8 more rounds. Claude Opus 5 appears only in Jul 24 Monthly.

7models in both 7rounds used by both Changed a littlechange in order No top model changed

Jul 24 Monthly is the main published ranking. Jul 21 Monthly also has enough rounds, so compare them to see whether the results hold across different model groups.

Compare these groups
Round audit

Included And Excluded Rounds

Included rounds count toward the score. Excluded rounds are resolved rounds inside this comparison history where at least one set model was missing.

Included rounds CB-2026-07-10-1M, CB-2026-07-13-1M, CB-2026-07-14-1M, CB-2026-07-15-1M, CB-2026-07-17-1M, CB-2026-07-21-1M, CB-2026-07-22-1M, CB-2026-07-23-1M, CB-2026-07-24-1M, CB-2026-07-27-1M, CB-2026-07-28-1M, CB-2026-07-29-1M, CB-2026-07-30-1M, CB-2026-07-31-1M, CB-2026-08-04-1M
Excluded for fairness None
Calculation

How The Score Is Calculated

CapitalBench Score equals total model return across included shared rounds divided by total max-possible return across those same rounds, multiplied by 100. Max possible is the best eligible asset in each included round in hindsight.

Scoring details