Results

Benchmark Results

AI model portfolios are scored in separate weekly and monthly tracks using the same rules, frozen portfolios, and real public-market prices.

Benchmark status
129 completed 15 live

Weekly and monthly tracks are scored separately.

Current weekly benchmark leader
Grok 4.5 20.6 CapitalBench Score
Latest scored CB-2026-08-25-1M Latest live CB-2026-09-21-1W / CB-2026-09-21-1M Models 12 Universe 70 options
  1. Completed
  2. Live
  3. Scored
Results insights

What Resolved Results Reveal

Signals generated from scored rounds, market environments, oracle comparisons, benchmark difficulty, and model confidence behavior.

Market EnvironmentAs of Sep 25
Monthly market environments12 resolved rounds2 models

Monthly model leadership changes with the S&P 500 environment

Grok 4.3 leads down environments at -1.22% across 6 tests; Grok 4.5 leads up environments at +4.11% across 6 tests.

Market environments group resolved rounds by the S&P 500 return over the same weekly or monthly window. Models are compared only on shared rounds; high confidence requires at least six observations and stable leadership.

Medium confidenceMath: deterministicData through Sep 25, 2026
Down Leader Average Return
-1.22%
Down Shared Rounds
6
Down Leader Stability
1.00
Consensus PerformanceAug 25-Sep 25
Monthly resultCB-2026-08-25-1M7 models

AI consensus portfolio scored -23.4 versus the oracle

If the monthly model allocations were averaged into one consensus portfolio, it returned -3.86% versus +0.69% for the S&P 500 and +16.47% for the hindsight best asset.

Consensus means the average of model allocations in the same round. CapitalBench Score compares that return with the hindsight-best eligible asset for that exact scoring window.

High confidenceMath: deterministicData through Sep 25, 2026
Consensus Portfolio Return
-3.86%
Average Model Return
-3.86%
Consensus Capitalbench Score
-23.4
Why it matters

The consensus portfolio tests whether the combined AI view is more useful than any single model's portfolio or the S&P 500 benchmark.

Benchmark DifficultyAug 25-Sep 25
Monthly resultCB-2026-08-25-1M7 models

Monthly round had +26.43% asset dispersion

The best scored asset returned +16.47%, the worst returned -9.96%, and +32.86% of the universe was positive. The S&P 500 ranked 21 out of 70 options.

Asset dispersion is the gap between the best and worst eligible assets in the same round. Wider dispersion makes missed allocation choices more costly.

High confidenceMath: deterministicData through Sep 25, 2026
Oracle Return
+16.5%
Worst Asset Return
-9.96%
Positive Universe Share
+32.9%
Why it matters

Benchmark difficulty matters because model scores should be interpreted against the opportunity set and the market window they faced.

Weekly track

Grok 4.5 Leads

CapitalBench Score leader inside the featured equal-run comparison set.

7 scored
Grok 4.5 xAI
CapitalBench Score 20.6 Avg return leader 2.02% Shared rounds 7 Timeline One market week
Monthly track

Claude Fable 5 Leads

CapitalBench Score leader inside the featured equal-run comparison set.

9 scored
Claude Fable 5 Anthropic
CapitalBench Score 8.4 Avg return leader 1.94% Shared rounds 9 Timeline One market month
Completed results

Current Benchmark Scores

These are equal-run comparison sets. Every ranked model completed every included round, and missed rounds are excluded from the set for everyone.

Open comparison sets
Equal-run benchmark

Current Weekly Benchmark

Every ranked model in this set completed the same 7 weekly rounds.

7 shared resolved rounds7 equal-run models rankedQualified at 6+ shared roundsNewest included round: CB-2026-09-16-1W
Shared resolved rounds

CapitalBench Score

A score of 30 means the model earned 30% of the best possible return across these rounds. Calculation

Grok 4.5
Grok 4.6
Gemini 3.1 Pro
Claude Opus 5
Claude Fable 5.1
GPT-6 Astra
Grok 4.3
S&P 500
Max possible What is this? Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio. hindsight best asset

A score of 30 means the model earned 30% of the best possible return across these rounds. Calculation

Grok 4.5 xAI · 7/7 scored rounds
20.6
Grok 4.6 xAI · 7/7 scored rounds
17.9
Gemini 3.1 Pro Google · 7/7 scored rounds
16.6
Claude Opus 5 Anthropic · 7/7 scored rounds
14.2
Claude Fable 5.1 Anthropic · 7/7 scored rounds
7.6
GPT-6 Astra OpenAI · 7/7 scored rounds
3.9
Grok 4.3 xAI · 7/7 scored rounds
-1.4
S&P 500 S&P 500 · 7/7 scored rounds
2.8
Max possible Hindsight ceiling, not a model portfolio
What is this? Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio. 100.0
7 shared resolved rounds7 equal-run models rankedQualified at 6+ shared roundsNewest included round: CB-2026-09-16-1W
Return context

Average Return Details

Average portfolio return across the same finished rounds.

xAI Grok 4.5
2.02%
xAI Grok 4.6
1.75%
Google Gemini 3.1 Pro
1.62%
Anthropic Claude Opus 5
1.39%
Anthropic Claude Fable 5.1
0.75%
OpenAI GPT-6 Astra
0.38%
xAI Grok 4.3
-0.14%
S&P S&P 500
0.27%
MAX Max possible What is this? Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio.
9.78%
Fairness rule: every ranked model completed every included round. A missed round is excluded from this set for everyone.
Equal-run benchmark

Current Monthly Benchmark

Every ranked model in this set completed the same 9 monthly rounds.

9 shared resolved rounds7 equal-run models rankedQualified at 3+ shared roundsNewest included round: CB-2026-08-25-1M
Shared resolved rounds

CapitalBench Score

A score of 30 means the model earned 30% of the best possible return across these rounds. Calculation

Claude Fable 5
GPT-5.6 Sol
Claude Opus 5
Grok 4.3
Grok 4.5
Grok 4.6
Gemini 3.1 Pro
S&P 500
Max possible What is this? Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio. hindsight best asset

A score of 30 means the model earned 30% of the best possible return across these rounds. Calculation

Claude Fable 5 Anthropic · 9/9 scored rounds
8.4
GPT-5.6 Sol OpenAI · 9/9 scored rounds
2.8
Claude Opus 5 Anthropic · 9/9 scored rounds
0.3
Grok 4.3 xAI · 9/9 scored rounds
-0.2
Grok 4.5 xAI · 9/9 scored rounds
-2.1
Grok 4.6 xAI · 9/9 scored rounds
-4.3
Gemini 3.1 Pro Google · 9/9 scored rounds
-6.8
S&P 500 S&P 500 · 9/9 scored rounds
-0.4
Max possible Hindsight ceiling, not a model portfolio
What is this? Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio. 100.0
9 shared resolved rounds7 equal-run models rankedQualified at 3+ shared roundsNewest included round: CB-2026-08-25-1M
Return context

Average Return Details

Average portfolio return across the same finished rounds.

Anthropic Claude Fable 5
1.94%
OpenAI GPT-5.6 Sol
0.65%
Anthropic Claude Opus 5
0.07%
xAI Grok 4.3
-0.05%
xAI Grok 4.5
-0.49%
xAI Grok 4.6
-0.99%
Google Gemini 3.1 Pro
-1.59%
S&P S&P 500
-0.10%
MAX Max possible What is this? Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio.
23.28%
Fairness rule: every ranked model completed every included round. A missed round is excluded from this set for everyone.
Fair model comparisons

Benchmark Comparison Sets

Sets are living groups. Older sets keep adding shared rounds, while newer model rosters become current automatically after enough shared results.

View all sets
Forming set Monthly Set: Sep 4, 2026

0 shared resolved rounds across 7 models.

3 more shared rounds to qualify
Forming set Monthly Set: Sep 3, 2026

0 shared resolved rounds across 7 models.

3 more shared rounds to qualify
Current benchmark Monthly Set: Aug 19, 2026

9 shared resolved rounds across 7 models.

Qualified at 3+ shared rounds
Qualified set Monthly Set: Aug 13, 2026

3 shared resolved rounds across 8 models.

Qualified at 3+ shared rounds
Qualified set Monthly Set: Jul 24, 2026

11 shared resolved rounds across 8 models.

Qualified at 3+ shared rounds
Qualified set Monthly Set: Jul 21, 2026

19 shared resolved rounds across 7 models.

Qualified at 3+ shared rounds
Qualified set Monthly Set: Jul 10, 2026

5 shared resolved rounds across 8 models.

Qualified at 3+ shared rounds
Qualified set Monthly Set: Jul 8, 2026

7 shared resolved rounds across 7 models.

Qualified at 3+ shared rounds
Qualified set Monthly Set: Jun 9, 2026

13 shared resolved rounds across 6 models.

Qualified at 3+ shared rounds
Qualified set Monthly Set: May 28, 2026

32 shared resolved rounds across 5 models.

Qualified at 3+ shared rounds
Qualified set Monthly Set: May 10, 2026

35 shared resolved rounds across 4 models.

Qualified at 3+ shared rounds
Current benchmark Weekly Set: Sep 4, 2026

7 shared resolved rounds across 7 models.

Qualified at 6+ shared rounds
Forming set Weekly Set: Sep 3, 2026

1 shared resolved rounds across 7 models.

5 more shared rounds to qualify
Qualified set Weekly Set: Aug 19, 2026

14 shared resolved rounds across 7 models.

Qualified at 6+ shared rounds
Forming set Weekly Set: Aug 13, 2026

3 shared resolved rounds across 8 models.

3 more shared rounds to qualify
Qualified set Weekly Set: Jul 24, 2026

11 shared resolved rounds across 8 models.

Qualified at 6+ shared rounds
Qualified set Weekly Set: Jul 21, 2026

20 shared resolved rounds across 7 models.

Qualified at 6+ shared rounds
Qualified set Weekly Set: Jul 10, 2026

6 shared resolved rounds across 8 models.

Qualified at 6+ shared rounds
Qualified set Weekly Set: Jul 8, 2026

8 shared resolved rounds across 7 models.

Qualified at 6+ shared rounds
Qualified set Weekly Set: Jun 9, 2026

15 shared resolved rounds across 6 models.

Qualified at 6+ shared rounds
Qualified set Weekly Set: May 28, 2026

33 shared resolved rounds across 5 models.

Qualified at 6+ shared rounds
Qualified set Weekly Set: May 24, 2026

35 shared resolved rounds across 4 models.

Qualified at 6+ shared rounds
Latest scored round

Most Recent Published Result

This chart shows the newest completed round only. Live rounds stay out of this chart until ending prices are collected.

Monthly result

Monthly Portfolio Returns

Models, S&P 500, and maximum possible return are shown on one scale.

Scored
Model portfolios S&P 500 benchmark Maximum possible return
GPT-5.6 Sol
Claude Fable 5
Grok 4.3
Claude Opus 5
Grok 4.6
Grok 4.5
Gemini 3.1 Pro
S&P 500
USO Crude Oil
GPT-5.6 Sol OpenAI
3.57%
Claude Fable 5 Anthropic
0.13%
Grok 4.3 xAI
-3.70%
Claude Opus 5 Anthropic
-4.80%
Grok 4.6 xAI
-6.35%
Grok 4.5 xAI
-6.41%
Gemini 3.1 Pro Google
-9.47%
S&P 500 Benchmark
0.69%
USO Crude Oil - Hindsight best asset
16.47%
Portfolio context

Shows each model's saved portfolio weights.

Model portfolios

Ranked in the same order as the chart.

1
GPT-5.6 Sol OpenAI
Semiconductors (SMH) 35% Cybersecurity (CIBR) 35% Real Estate (XLRE) 30%
2
Claude Fable 5 Anthropic
Semiconductors (SMH) 35% Regional Banks (KRE) 35% Industrials (XLI) 30%
3
Grok 4.3 xAI
US Dollar (UUP) 35% Small Value (IWN) 35% Utilities (XLU) 30%
4
Claude Opus 5 Anthropic
Industrials (XLI) 35% Regional Banks (KRE) 35% Small Value (IWN) 30%
5
Grok 4.6 xAI
Defense (ITA) 35% Utilities (XLU) 35% S&P 500 (SPY) 30%
6
Grok 4.5 xAI
Defense (ITA) 35% Regional Banks (KRE) 35% Industrials (XLI) 30%
7
Gemini 3.1 Pro Google
Solar (TAN) 35% Utilities (XLU) 35% Defense (ITA) 30%
Reference points

Not model portfolios.

S&P 500 Benchmark

Benchmark return over the same scoring window

USO Crude Oil - Hindsight best asset

100% Crude Oil (USO) hindsight ceiling

Run details

CB-2026-08-25-1M

2026-08-26 to 2026-09-25

WinnerGPT-5.6 Sol Return3.57% Models7 Eligible assets70
Benchmark result paths

Choose The Result View

Latest pages show one completed round. All-history pages are context. Comparison sets are the fair ranking view.

All rounds
Weekly result Latest Weekly

CB-2026-09-16-1W, 2026-09-17 to 2026-09-24

Scored
Weekly aggregate Overall Weekly

71 all-available weekly rounds for context. Fair rankings use comparison sets.

7 rounds
Monthly result Latest Monthly

CB-2026-08-25-1M, 2026-08-26 to 2026-09-25

Scored
Monthly aggregate Overall Monthly

58 all-available monthly rounds for context. Fair rankings use comparison sets.

9 rounds
Live rounds

Waiting For Final Prices

These rounds are live or pending score. They are not counted in completed result charts yet.

Weekly live round

CB-2026-09-21-1W

Scores after the 2026-09-28 close.

2026-09-21 to 2026-09-28 official-v3-20260921-weekly
View locked portfolios
Monthly live round

CB-2026-09-21-1M

Scores after the 2026-10-21 close.

2026-09-21 to 2026-10-21 official-v3-20260921-monthly
View locked portfolios
Audit and data

Every Result Has An Audit Packet

Results link back to prompts, model outputs, portfolio decisions, prices, audit hashes, and scoring records.