Results

Benchmark Results

AI model portfolios are scored in separate weekly and monthly tracks using the same rules, frozen portfolios, and real public-market prices.

Benchmark status
77 completed 23 live

Weekly and monthly tracks are scored separately.

Current weekly benchmark leader
Grok 4.3 11.5 CapitalBench Score
Latest scored CB-2026-08-05-1W Latest live CB-2026-08-13-1W / CB-2026-08-13-1M Models 10 Universe 70 options
  1. Completed
  2. Live
  3. Scored
Results insights

What Resolved Results Reveal

Signals generated from scored rounds, market environments, oracle comparisons, benchmark difficulty, and model confidence behavior.

Market EnvironmentAs of Aug 10
Monthly market environments9 resolved rounds1 model

Grok 4.3 leads across multiple monthly market environments

Grok 4.3 leads down environments at -1.22% across 6 tests; Grok 4.3 leads up environments at +2.26% across 3 tests.

Market environments group resolved rounds by the S&P 500 return over the same weekly or monthly window. Models are compared only on shared rounds; high confidence requires at least six observations and stable leadership.

Medium confidenceMath: deterministicData through Aug 10, 2026
Down Leader Average Return
-1.22%
Down Shared Rounds
6
Down Leader Stability
1.00
Confidence CalibrationAs of Aug 13
All resolved official results77 resolved rounds452 scored results

High-confidence model calls have underperformed lower-confidence calls

Across resolved official results, submissions at or above the median confidence of 0.58 averaged -1.18%, while lower-confidence submissions averaged -0.20%.

Confidence is the model's own 0-1 self-reported confidence at submission time, compared with later realized returns.

High confidenceMath: deterministicData through Aug 13, 2026
High Confidence Average Return
-1.18%
Low Confidence Average Return
-0.20%
High Confidence Average Capitalbench Score
-11.4
Why it matters

Confidence calibration helps readers judge whether model self-reported confidence carries useful information about realized benchmark performance.

Consensus PerformanceAug 5-Aug 12
Weekly resultCB-2026-08-05-1W8 models

AI consensus portfolio scored 12.5 versus the oracle

If the weekly model allocations were averaged into one consensus portfolio, it returned +1.35% versus +0.35% for the S&P 500 and +10.81% for the hindsight best asset.

Consensus means the average of model allocations in the same round. CapitalBench Score compares that return with the hindsight-best eligible asset for that exact scoring window.

High confidenceMath: deterministicData through Aug 12, 2026
Consensus Portfolio Return
+1.35%
Average Model Return
+1.35%
Consensus Capitalbench Score
12.5
Why it matters

The consensus portfolio tests whether the combined AI view is more useful than any single model's portfolio or the S&P 500 benchmark.

Weekly track

Grok 4.3 Leads

CapitalBench Score leader inside the featured equal-run comparison set.

8 scored
Grok 4.3 xAI
CapitalBench Score 11.5 Avg return leader 1.29% Shared rounds 8 Timeline One market week
Monthly track

Grok 4.3 Leads

CapitalBench Score leader inside the featured equal-run comparison set.

3 scored
Grok 4.3 xAI
CapitalBench Score 16.3 Avg return leader 2.26% Shared rounds 3 Timeline One market month
Completed results

Current Benchmark Scores

These are equal-run comparison sets. Every ranked model completed every included round, and missed rounds are excluded from the set for everyone.

Open comparison sets
Equal-run benchmark

Current Weekly Benchmark

Every ranked model in this set completed the same 8 weekly rounds.

8 shared resolved rounds8 equal-run models rankedQualified at 6+ shared roundsNewest included round: CB-2026-08-05-1W
Shared resolved rounds

CapitalBench Score

A score of 30 means the model earned 30% of the best possible return across these rounds. Calculation

Grok 4.3
Grok 4.5
Claude Opus 5
GPT-5.6 Sol
Claude Opus 4.8
Claude Fable 5
GPT-5.5
Gemini 3.1 Pro
S&P 500
Max possible What is this? Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio. hindsight best asset

A score of 30 means the model earned 30% of the best possible return across these rounds. Calculation

Grok 4.3 xAI · 8/8 scored rounds
11.5
Grok 4.5 xAI · 8/8 scored rounds
10.8
Claude Opus 5 Anthropic · 8/8 scored rounds
9.8
GPT-5.6 Sol OpenAI · 8/8 scored rounds
8.8
Claude Opus 4.8 Anthropic · 8/8 scored rounds
7.2
Claude Fable 5 Anthropic · 8/8 scored rounds
6.2
GPT-5.5 OpenAI · 8/8 scored rounds
4.4
Gemini 3.1 Pro Google · 8/8 scored rounds
2.3
S&P 500 S&P 500 · 8/8 scored rounds
23.1
Max possible Hindsight ceiling, not a model portfolio
What is this? Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio. 100.0
8 shared resolved rounds8 equal-run models rankedQualified at 6+ shared roundsNewest included round: CB-2026-08-05-1W
Return context

Average Return Details

Average portfolio return across the same finished rounds.

xAI Grok 4.3
1.29%
xAI Grok 4.5
1.20%
Anthropic Claude Opus 5
1.10%
OpenAI GPT-5.6 Sol
0.98%
Anthropic Claude Opus 4.8
0.80%
Anthropic Claude Fable 5
0.69%
OpenAI GPT-5.5
0.50%
Google Gemini 3.1 Pro
0.26%
S&P S&P 500
2.58%
MAX Max possible What is this? Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio.
11.18%
Fairness rule: every ranked model completed every included round. A missed round is excluded from this set for everyone.
Equal-run benchmark

Current Monthly Benchmark

Every ranked model in this set completed the same 3 monthly rounds.

3 shared resolved rounds7 equal-run models rankedQualified at 3+ shared roundsNewest included round: CB-2026-07-10-1M
Shared resolved rounds

CapitalBench Score

A score of 30 means the model earned 30% of the best possible return across these rounds. Calculation

Grok 4.3
Claude Opus 4.7
Claude Opus 4.8
Claude Fable 5
GPT-5.5
Gemini 3.1 Pro
Grok 4.5
S&P 500
Max possible What is this? Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio. hindsight best asset

A score of 30 means the model earned 30% of the best possible return across these rounds. Calculation

Grok 4.3 xAI · 3/3 scored rounds
16.3
Claude Opus 4.7 Anthropic · 3/3 scored rounds
7.8
Claude Opus 4.8 Anthropic · 3/3 scored rounds
7.8
Claude Fable 5 Anthropic · 3/3 scored rounds
1.1
GPT-5.5 OpenAI · 3/3 scored rounds
-2.3
Gemini 3.1 Pro Google · 3/3 scored rounds
-8.0
Grok 4.5 xAI · 3/3 scored rounds
-8.9
S&P 500 S&P 500 · 3/3 scored rounds
21.7
Max possible Hindsight ceiling, not a model portfolio
What is this? Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio. 100.0
3 shared resolved rounds7 equal-run models rankedQualified at 3+ shared roundsNewest included round: CB-2026-07-10-1M
Return context

Average Return Details

Average portfolio return across the same finished rounds.

xAI Grok 4.3
2.26%
Anthropic Claude Opus 4.7
1.08%
Anthropic Claude Opus 4.8
1.07%
Anthropic Claude Fable 5
0.16%
OpenAI GPT-5.5
-0.32%
Google Gemini 3.1 Pro
-1.11%
xAI Grok 4.5
-1.22%
S&P S&P 500
3.00%
MAX Max possible What is this? Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio.
13.83%
Fairness rule: every ranked model completed every included round. A missed round is excluded from this set for everyone.
Fair model comparisons

Benchmark Comparison Sets

Sets are living groups. Older sets keep adding shared rounds, while newer model rosters become current automatically after enough shared results.

View all sets
Forming set Monthly Set: Aug 13, 2026

0 shared resolved rounds across 8 models.

3 more shared rounds to qualify
Forming set Monthly Set: Jul 24, 2026

0 shared resolved rounds across 8 models.

3 more shared rounds to qualify
Forming set Monthly Set: Jul 21, 2026

0 shared resolved rounds across 7 models.

3 more shared rounds to qualify
Forming set Monthly Set: Jul 10, 2026

1 shared resolved rounds across 8 models.

2 more shared rounds to qualify
Current benchmark Monthly Set: Jul 8, 2026

3 shared resolved rounds across 7 models.

Qualified at 3+ shared rounds
Qualified set Monthly Set: Jun 9, 2026

9 shared resolved rounds across 6 models.

Qualified at 3+ shared rounds
Qualified set Monthly Set: May 28, 2026

28 shared resolved rounds across 5 models.

Qualified at 3+ shared rounds
Qualified set Monthly Set: May 10, 2026

31 shared resolved rounds across 4 models.

Qualified at 3+ shared rounds
Forming set Weekly Set: Aug 13, 2026

0 shared resolved rounds across 8 models.

6 more shared rounds to qualify
Current benchmark Weekly Set: Jul 24, 2026

8 shared resolved rounds across 8 models.

Qualified at 6+ shared rounds
Qualified set Weekly Set: Jul 21, 2026

11 shared resolved rounds across 7 models.

Qualified at 6+ shared rounds
Qualified set Weekly Set: Jul 10, 2026

6 shared resolved rounds across 8 models.

Qualified at 6+ shared rounds
Qualified set Weekly Set: Jul 8, 2026

8 shared resolved rounds across 7 models.

Qualified at 6+ shared rounds
Qualified set Weekly Set: Jun 9, 2026

15 shared resolved rounds across 6 models.

Qualified at 6+ shared rounds
Qualified set Weekly Set: May 28, 2026

33 shared resolved rounds across 5 models.

Qualified at 6+ shared rounds
Qualified set Weekly Set: May 24, 2026

35 shared resolved rounds across 4 models.

Qualified at 6+ shared rounds
Latest scored round

Most Recent Published Result

This chart shows the newest completed round only. Live rounds stay out of this chart until ending prices are collected.

Weekly result

Weekly Portfolio Returns

Models, S&P 500, and maximum possible return are shown on one scale.

Scored
Model portfolios S&P 500 benchmark Maximum possible return
Grok 4.5
GPT-5.6 Sol
Claude Opus 5
GPT-5.5
Claude Fable 5
Claude Opus 4.8
Grok 4.3
Gemini 3.1 Pro
S&P 500
USO Crude Oil
Grok 4.5 xAI
2.93%
GPT-5.6 Sol OpenAI
2.54%
Claude Opus 5 Anthropic
1.42%
GPT-5.5 OpenAI
1.40%
Claude Fable 5 Anthropic
1.15%
Claude Opus 4.8 Anthropic
0.68%
Grok 4.3 xAI
0.35%
Gemini 3.1 Pro Google
0.35%
S&P 500 Benchmark
0.35%
USO Crude Oil - Hindsight best asset
10.81%
Portfolio context

Shows each model's saved portfolio weights.

Model portfolios

Ranked in the same order as the chart.

1
Grok 4.5 xAI
Technology (XLK) 30% Energy (XLE) 25% Gold (IAU) 20% Financials (XLF) 15% Value (IWD) 10%
2
GPT-5.6 Sol OpenAI
Dividend (SCHD) 30% Energy (XLE) 30% Real Estate (XLRE) 20% Software (IGV) 20%
3
Claude Opus 5 Anthropic
S&P 500 (SPY) 40% Healthcare (XLV) 20% Gold (IAU) 20% Financials (XLF) 20%
4
GPT-5.5 OpenAI
Energy (XLE) 25% Real Estate (XLRE) 25% Software (IGV) 25% China (MCHI) 15% Financials (XLF) 10%
5
Claude Fable 5 Anthropic
S&P 500 (SPY) 35% Dividend (SCHD) 30% Healthcare (XLV) 20% Staples (XLP) 15%
6
Claude Opus 4.8 Anthropic
S&P 500 (SPY) 50% Technology (XLK) 25% Growth (IWF) 25%
7
Grok 4.3 xAI
S&P 500 (SPY) 100%
8
Gemini 3.1 Pro Google
S&P 500 (SPY) 100%
Reference points

Not model portfolios.

S&P 500 Benchmark

Benchmark return over the same scoring window

USO Crude Oil - Hindsight best asset

100% Crude Oil (USO) hindsight ceiling

Run details

CB-2026-08-05-1W

2026-08-05 to 2026-08-12

WinnerGrok 4.5 Return2.93% Models8 Eligible assets70
Benchmark result paths

Choose The Result View

Latest pages show one completed round. All-history pages are context. Comparison sets are the fair ranking view.

All rounds
Weekly result Latest Weekly

CB-2026-08-05-1W, 2026-08-05 to 2026-08-12

Scored
Weekly aggregate Overall Weekly

46 all-available weekly rounds for context. Fair rankings use comparison sets.

8 rounds
Monthly result Latest Monthly

CB-2026-07-10-1M, 2026-07-10 to 2026-08-10

Scored
Monthly aggregate Overall Monthly

31 all-available monthly rounds for context. Fair rankings use comparison sets.

3 rounds
Live rounds

Waiting For Final Prices

These rounds are live or pending score. They are not counted in completed result charts yet.

Weekly live round

CB-2026-08-13-1W

Scores after the 2026-08-20 close.

2026-08-13 to 2026-08-20 official-v2-2-all-weekly-20260813
View locked portfolios
Monthly live round

CB-2026-08-13-1M

Scores after the 2026-09-14 close.

2026-08-13 to 2026-09-14 official-v2-2-all-monthly-20260813
View locked portfolios
Audit and data

Every Result Has An Audit Packet

Results link back to prompts, model outputs, portfolio decisions, prices, audit hashes, and scoring records.