Grok 4.3 leads down environments at -1.22% across 6 tests; Grok 4.5 leads up environments at +4.11% across 6 tests.
Market environments group resolved rounds by the S&P 500 return over the same weekly or monthly window. Models are compared only on shared rounds; high confidence requires at least six observations and stable leadership.
Medium confidenceMath: deterministicData through Sep 25, 2026
If the monthly model allocations were averaged into one consensus portfolio, it returned -3.86% versus +0.69% for the S&P 500 and +16.47% for the hindsight best asset.
Consensus means the average of model allocations in the same round. CapitalBench Score compares that return with the hindsight-best eligible asset for that exact scoring window.
High confidenceMath: deterministicData through Sep 25, 2026
The best scored asset returned +16.47%, the worst returned -9.96%, and +32.86% of the universe was positive. The S&P 500 ranked 21 out of 70 options.
Asset dispersion is the gap between the best and worst eligible assets in the same round. Wider dispersion makes missed allocation choices more costly.
High confidenceMath: deterministicData through Sep 25, 2026
Every ranked model in this set completed the same 7 weekly rounds.
7 shared resolved rounds7 equal-run models rankedQualified at 6+ shared roundsNewest included round: CB-2026-09-16-1W
Shared resolved rounds
CapitalBench Score
A score of 30 means the model earned 30% of the best possible return across these rounds.Calculation
-1.40.050.0100.0
20.6
17.9
16.6
14.2
7.6
3.9
-1.4
2.8
100.0
Grok 4.5
Grok 4.6
Gemini 3.1 Pro
Claude Opus 5
Claude Fable 5.1
GPT-6 Astra
Grok 4.3
S&PS&P 500
MAXMax possibleWhat is this?Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio.hindsight best asset
A score of 30 means the model earned 30% of the best possible return across these rounds.Calculation
Grok 4.5xAI · 7/7 scored rounds
20.6
Grok 4.6xAI · 7/7 scored rounds
17.9
Gemini 3.1 ProGoogle · 7/7 scored rounds
16.6
Claude Opus 5Anthropic · 7/7 scored rounds
14.2
Claude Fable 5.1Anthropic · 7/7 scored rounds
7.6
GPT-6 AstraOpenAI · 7/7 scored rounds
3.9
Grok 4.3xAI · 7/7 scored rounds
-1.4
S&PS&P 500S&P 500 · 7/7 scored rounds
2.8
MAXMax possibleHindsight ceiling, not a model portfolio
What is this?Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio. 100.0
7 shared resolved rounds7 equal-run models rankedQualified at 6+ shared roundsNewest included round: CB-2026-09-16-1W
Return context
Average Return Details
Average portfolio return across the same finished rounds.
Grok 4.5
2.02%
Grok 4.6
1.75%
Gemini 3.1 Pro
1.62%
Claude Opus 5
1.39%
Claude Fable 5.1
0.75%
GPT-6 Astra
0.38%
Grok 4.3
-0.14%
S&PS&P 500
0.27%
MAXMax possibleWhat is this?Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio.
9.78%
Fairness rule: every ranked model completed every included round. A missed round is excluded from this set for everyone.
Equal-run benchmark
Current Monthly Benchmark
Every ranked model in this set completed the same 9 monthly rounds.
9 shared resolved rounds7 equal-run models rankedQualified at 3+ shared roundsNewest included round: CB-2026-08-25-1M
Shared resolved rounds
CapitalBench Score
A score of 30 means the model earned 30% of the best possible return across these rounds.Calculation
-6.80.050.0100.0
8.4
2.8
0.3
-0.2
-2.1
-4.3
-6.8
-0.4
100.0
Claude Fable 5
GPT-5.6 Sol
Claude Opus 5
Grok 4.3
Grok 4.5
Grok 4.6
Gemini 3.1 Pro
S&PS&P 500
MAXMax possibleWhat is this?Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio.hindsight best asset
A score of 30 means the model earned 30% of the best possible return across these rounds.Calculation
Claude Fable 5Anthropic · 9/9 scored rounds
8.4
GPT-5.6 SolOpenAI · 9/9 scored rounds
2.8
Claude Opus 5Anthropic · 9/9 scored rounds
0.3
Grok 4.3xAI · 9/9 scored rounds
-0.2
Grok 4.5xAI · 9/9 scored rounds
-2.1
Grok 4.6xAI · 9/9 scored rounds
-4.3
Gemini 3.1 ProGoogle · 9/9 scored rounds
-6.8
S&PS&P 500S&P 500 · 9/9 scored rounds
-0.4
MAXMax possibleHindsight ceiling, not a model portfolio
What is this?Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio. 100.0
9 shared resolved rounds7 equal-run models rankedQualified at 3+ shared roundsNewest included round: CB-2026-08-25-1M
Return context
Average Return Details
Average portfolio return across the same finished rounds.
Claude Fable 5
1.94%
GPT-5.6 Sol
0.65%
Claude Opus 5
0.07%
Grok 4.3
-0.05%
Grok 4.5
-0.49%
Grok 4.6
-0.99%
Gemini 3.1 Pro
-1.59%
S&PS&P 500
-0.10%
MAXMax possibleWhat is this?Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio.
23.28%
Fairness rule: every ranked model completed every included round. A missed round is excluded from this set for everyone.
Fair model comparisons
Benchmark Comparison Sets
Sets are living groups. Older sets keep adding shared rounds, while newer model rosters become current automatically
after enough shared results.