CapitalBench

The benchmark for AI capital allocation

Each AI model gets the same market brief, builds one portfolio in one response, and uses no browsing or tools. Real market prices determine the result.

See how AI models perform against each other, how they invest and take risk, and how they perform in the real market.

Read the CapitalBench Manifesto
Benchmark results

Which models are performing best?

Monthly and weekly tracks stay separate.

Current Monthly Benchmark

CapitalBench Score

A score of 30 means the model earned 30% of the best possible return across these rounds. Calculation

Claude Fable 5
GPT-5.6 Sol
Claude Opus 5
Grok 4.3
Grok 4.5
Grok 4.6
Gemini 3.1 Pro
S&P 500
Max possible What is this? Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio. hindsight best asset

A score of 30 means the model earned 30% of the best possible return across these rounds. Calculation

Claude Fable 5 Anthropic · 9/9 scored rounds
8.4
GPT-5.6 Sol OpenAI · 9/9 scored rounds
2.8
Claude Opus 5 Anthropic · 9/9 scored rounds
0.3
Grok 4.3 xAI · 9/9 scored rounds
-0.2
Grok 4.5 xAI · 9/9 scored rounds
-2.1
Grok 4.6 xAI · 9/9 scored rounds
-4.3
Gemini 3.1 Pro Google · 9/9 scored rounds
-6.8
S&P 500 S&P 500 · 9/9 scored rounds
-0.4
Max possible Hindsight ceiling, not a model portfolio
What is this? Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio. 100.0
9 shared resolved rounds7 equal-run models rankedQualified at 3+ shared roundsNewest included round: CB-2026-08-25-1M
Return context

Average Return Details

Average portfolio return across the same finished rounds.

Anthropic Claude Fable 5
1.94%
OpenAI GPT-5.6 Sol
0.65%
Anthropic Claude Opus 5
0.07%
xAI Grok 4.3
-0.05%
xAI Grok 4.5
-0.49%
xAI Grok 4.6
-0.99%
Google Gemini 3.1 Pro
-1.59%
S&P S&P 500
-0.10%
MAX Max possible What is this? Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio.
23.28%
Fairness rule: every ranked model completed every included round. A missed round is excluded from this set for everyone.

Monthly and weekly are separate comparison tracks. Scores are never mixed across horizons. Read scoring rules

Model Risk Benchmark

Which AI models take the most risk?

Compare how aggressively each model allocates capital across its official portfolios.

Updated through September 21, 2026Higher means more risk-seeking, not better.
GPT-5.5OpenAI
79.7/100Risk-seeking
98portfolios34.1%largest position3.8%defensive assets
Claude Fable 5.1Anthropic
75.1/100Risk-seeking
18portfolios40.0%largest position7.5%defensive assets
GPT-5.6 SolOpenAI
73.6/100Risk-seeking
69portfolios36.1%largest position8.1%defensive assets
Grok 4.5xAI
73.4/100Risk-seeking
89portfolios34.8%largest position6.7%defensive assets
Grok 4.3xAI
72.7/100Risk-seeking
144portfolios49.2%largest position6.3%defensive assets
Claude Opus 5Anthropic
72.3/100Risk-seeking
68portfolios35.7%largest position7.5%defensive assets
Gemini 3.1 ProGoogle
71.4/100Risk-seeking
144portfolios39.9%largest position13.7%defensive assets
Claude Opus 4.7Anthropic
71.3/100Risk-seeking
70portfolios31.8%largest position16.6%defensive assets
Grok 4.6xAI
71.0/100Risk-seeking
46portfolios62.3%largest position5.5%defensive assets
Claude Fable 5Anthropic
70.8/100Risk-seeking
84portfolios31.3%largest position12.1%defensive assets
GPT-6 AstraOpenAI
69.7/100Risk-seeking
16portfolios44.7%largest position10.9%defensive assets
Claude Opus 4.8Anthropic
69.2/100Risk-seeking
99portfolios35.3%largest position11.7%defensive assets
Portfolio Difference

Which AI models invest most differently from the group?

Compare each model's portfolio with the choices made by the other models in the same rounds.

Updated through September 21, 2026Overall: 50% monthly / 50% weekly
Grok 4.6xAI
59.1/100
44same-round comparisonsEstablished sample
Grok 4.5xAI
50.9/100
44same-round comparisonsEstablished sample
GPT-6 AstraOpenAI
48.6/100
16same-round comparisonsEstablished sample

Different does not mean better or worse. The score compares portfolio outputs; it does not prove copying, influence, or intent.

Explore Portfolio Difference
Performance by market

Who leads when markets rise or fall?

Monthly leaders vs. the S&P 500.

Monthly snapshot Updated Sep 25 4 of 5 market types comparable
Market fell S&P < -1.0%
Leader Grok 4.3
Average return -1.22%
Versus S&P 500 -1.90%
+0.68 pts vs S&P 6 results · Some evidence
Market rose S&P > +1.0%
Leader Grok 4.5
Average return +4.11%
Versus S&P 500 +3.90%
+0.22 pts vs S&P 6 results · Some evidence
AI positioning

What are AI models doing right now?

Live allocations before the next official score.

As of September 21, 2026
Current risk appetite As of September 21, 2026 72.5/100 Risk-seeking / Broad risk seeking
Consensus allocation As of September 21, 2026 41.8% S&P 500 (SPY) average live weight
Risk shift As of September 21, 2026 -8.1 Change vs Sep 16 portfolios
Model agreement As of September 21, 2026 Tight 3.5 point dispersion
Current risk appetite 72.5/100 As of September 21, 2026 / Risk-seeking
As of September 21, 2026 Broad risk seeking Combined view of monthly and weekly model portfolios for this date.
Unscored portfolios 73.2/100 Separate read across every open portfolio before official scoring.
Largest current allocations
S&P 500 (SPY) 41.8% Biotechnology (XBI) 12.5% Japan Equities (EWJ) 10.0% Metals and Mining (XME) 5.0% South Korea Equities (EWY) 5.0% Mexico Equities (EWW) 4.6%
Regime mix
Broad and cyclical equity 53.6% International equity 19.6% Growth and technology 12.5% Real assets and inflation 9.3% Defensive equity 2.5%
Trust and proof

What makes each model test comparable?

Same brief. Same choices. One response. No tools. Real market prices.

  1. Step 1 Same report

    Every model reads the same market report.

  2. Step 2 Same choices

    Every model chooses from the same 70 assets.

  3. Step 3 One response, then locked

    No browsing, tools, or follow-up prompts. The model's submitted portfolio is frozen before results are known.

  4. Step 4 Fixed wait window

    The frozen portfolio sits untouched for 7 days or 1 month.

  5. Step 5 Prices score it

    Real ending prices decide which model did best.

Benchmark universe

What can models choose from?

The active roster, asset menu, horizons, and open rounds.

Models 12
Asset choices 70
Round lengths 2
Open rounds 15
Protocol Single-turn Non-agentic calls
Latest official results

What happened in the latest scored rounds?

Finished monthly and weekly rounds scored against real market returns.

Monthly result1 of 58
Monthly official result

Monthly result scored Sep 25

Same-window returns, ranked after final prices.

Scored
Model portfolios S&P 500 benchmark Maximum possible return
GPT-5.6 Sol
Claude Fable 5
Grok 4.3
Claude Opus 5
Grok 4.6
Grok 4.5
Gemini 3.1 Pro
S&P 500
USO Crude Oil
GPT-5.6 Sol OpenAI
3.57%
Claude Fable 5 Anthropic
0.13%
Grok 4.3 xAI
-3.70%
Claude Opus 5 Anthropic
-4.80%
Grok 4.6 xAI
-6.35%
Grok 4.5 xAI
-6.41%
Gemini 3.1 Pro Google
-9.47%
S&P 500 Benchmark
0.69%
USO Crude Oil - Hindsight best asset
16.47%
Portfolio context

Shows each model's saved portfolio weights.

Model portfolios

Ranked in the same order as the chart.

1
GPT-5.6 Sol OpenAI
Semiconductors (SMH) 35% Cybersecurity (CIBR) 35% Real Estate (XLRE) 30%
2
Claude Fable 5 Anthropic
Semiconductors (SMH) 35% Regional Banks (KRE) 35% Industrials (XLI) 30%
3
Grok 4.3 xAI
US Dollar (UUP) 35% Small Value (IWN) 35% Utilities (XLU) 30%
4
Claude Opus 5 Anthropic
Industrials (XLI) 35% Regional Banks (KRE) 35% Small Value (IWN) 30%
5
Grok 4.6 xAI
Defense (ITA) 35% Utilities (XLU) 35% S&P 500 (SPY) 30%
6
Grok 4.5 xAI
Defense (ITA) 35% Regional Banks (KRE) 35% Industrials (XLI) 30%
7
Gemini 3.1 Pro Google
Solar (TAN) 35% Utilities (XLU) 35% Defense (ITA) 30%
Reference points

Not model portfolios.

S&P 500 Benchmark

Benchmark return over the same scoring window

USO Crude Oil - Hindsight best asset

100% Crude Oil (USO) hindsight ceiling

Official scored round

Monthly result scored Sep 25

Audit ID: CB-2026-08-25-1M

ScoredSep 25WindowAug 26 to Sep 25Models7Asset choices70LeaderGPT-5.6 SolHorizonMonthly
Live dashboard

What is still in progress?

Open portfolios, interim returns, and upcoming score dates.

Open rounds 15 14 monthly / 1 weekly
Frozen portfolios 105 29 assets currently held
Latest close Sep 25 Live returns update before final scoring
Next score Sep 28 Official results publish after ending prices
Audit packet

How can you verify the benchmark?

Round packets expose the report, prompt, portfolios, prices, hashes, and result status.