CapitalBench

The benchmark for AI capital allocation

Each AI model gets the same market brief, builds one portfolio in one response, and uses no browsing or tools. Real market prices determine the result.

See how AI models perform against each other, how they invest and take risk, and how they perform in the real market.

Read the CapitalBench Manifesto
Benchmark results

Which models are performing best?

Monthly and weekly tracks stay separate.

Current Monthly Benchmark

CapitalBench Score

A score of 30 means the model earned 30% of the best possible return across these rounds. Calculation

Grok 4.3
Claude Opus 4.7
Claude Opus 4.8
Claude Fable 5
GPT-5.5
Gemini 3.1 Pro
Grok 4.5
S&P 500
Max possible What is this? Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio. hindsight best asset

A score of 30 means the model earned 30% of the best possible return across these rounds. Calculation

Grok 4.3 xAI · 3/3 scored rounds
16.3
Claude Opus 4.7 Anthropic · 3/3 scored rounds
7.8
Claude Opus 4.8 Anthropic · 3/3 scored rounds
7.8
Claude Fable 5 Anthropic · 3/3 scored rounds
1.1
GPT-5.5 OpenAI · 3/3 scored rounds
-2.3
Gemini 3.1 Pro Google · 3/3 scored rounds
-8.0
Grok 4.5 xAI · 3/3 scored rounds
-8.9
S&P 500 S&P 500 · 3/3 scored rounds
21.7
Max possible Hindsight ceiling, not a model portfolio
What is this? Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio. 100.0
3 shared resolved rounds7 equal-run models rankedQualified at 3+ shared roundsNewest included round: CB-2026-07-10-1M
Return context

Average Return Details

Average portfolio return across the same finished rounds.

xAI Grok 4.3
2.26%
Anthropic Claude Opus 4.7
1.08%
Anthropic Claude Opus 4.8
1.07%
Anthropic Claude Fable 5
0.16%
OpenAI GPT-5.5
-0.32%
Google Gemini 3.1 Pro
-1.11%
xAI Grok 4.5
-1.22%
S&P S&P 500
3.00%
MAX Max possible What is this? Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio.
13.83%
Fairness rule: every ranked model completed every included round. A missed round is excluded from this set for everyone.

Monthly and weekly are separate comparison tracks. Scores are never mixed across horizons. Read scoring rules

Model Risk Benchmark

Which AI models take the most risk?

Compare how aggressively each model allocates capital across its official portfolios.

Updated through August 13, 2026Higher means more risk-seeking, not better.
GPT-5.5OpenAI
79.7/100Risk-seeking
98portfolios34.1%largest position3.8%defensive assets
Grok 4.3xAI
74.3/100Risk-seeking
100portfolios44.3%largest position5.1%defensive assets
Grok 4.5xAI
71.8/100Risk-seeking
45portfolios31.8%largest position3.8%defensive assets
GPT-5.6 SolOpenAI
71.7/100Risk-seeking
41portfolios36.8%largest position8.2%defensive assets
Claude Opus 4.7Anthropic
71.3/100Risk-seeking
70portfolios31.8%largest position16.6%defensive assets
Gemini 3.1 ProGoogle
70.4/100Risk-seeking
100portfolios41.8%largest position13.0%defensive assets
Claude Opus 4.8Anthropic
69.2/100Risk-seeking
95portfolios32.5%largest position12.2%defensive assets
Claude Fable 5Anthropic
67.6/100Risk-seeking
58portfolios28.5%largest position13.6%defensive assets
Claude Opus 5Anthropic
64.4/100Risk-seeking
24portfolios35.6%largest position10.2%defensive assets
Grok 4.6xAIEarly sample
64.3/100Risk-seeking
2portfolios30.0%largest position7.5%defensive assets
Portfolio Difference

Which AI models invest most differently from the group?

Compare each model's portfolio with the choices made by the other models in the same rounds.

Updated through August 13, 2026Overall: 50% monthly / 50% weekly
Grok 4.3xAI
58.5/100
30same-round comparisonsEstablished sample
Grok 4.5xAI
55.3/100
30same-round comparisonsEstablished sample
Grok 4.6xAI
47.9/100
2same-round comparisonsEarly sample

Different does not mean better or worse. The score compares portfolio outputs; it does not prove copying, influence, or intent.

Explore Portfolio Difference
Performance by market

Who leads when markets rise or fall?

Monthly leaders vs. the S&P 500.

Monthly snapshot Updated Aug 12 3 of 5 market types comparable
Market fell S&P < -1.0%
Leader Grok 4.3
Average return -1.22%
Versus S&P 500 -1.90%
+0.68 pts vs S&P 6 results · Some evidence
Market rose S&P > +1.0%
Leader Grok 4.3
Average return +2.26%
Versus S&P 500 +3.00%
-0.74 pts vs S&P 3 results · Some evidence
AI positioning

What are AI models doing right now?

Live allocations before the next official score.

As of August 13, 2026
Current risk appetite As of August 13, 2026 61.0/100 Risk-seeking / Selective risk taking
Consensus allocation As of August 13, 2026 27.8% Healthcare Sector (XLV) average live weight
Risk shift As of August 13, 2026 -2.9 Change vs Aug 11 portfolios
Model agreement As of August 13, 2026 Tight 3.7 point dispersion
Current risk appetite 61.0/100 As of August 13, 2026 / Risk-seeking
As of August 13, 2026 Selective risk taking Combined view of monthly and weekly model portfolios for this date.
Unscored portfolios 64.3/100 Separate read across every open portfolio before official scoring.
Largest current allocations
Healthcare Sector (XLV) 27.8% S&P 500 (SPY) 17.2% Energy Sector (XLE) 14.7% Gold (IAU) 10.9% Equal-Weight S&P 500 (RSP) 5.3% Japan Equities (EWJ) 4.4%
Regime mix
Defensive equity 31.9% Broad and cyclical equity 28.8% Real assets and inflation 26.6% Growth and technology 7.2% International equity 5.6%
Trust and proof

What makes each model test comparable?

Same brief. Same choices. One response. No tools. Real market prices.

  1. Step 1 Same report

    Every model reads the same market report.

  2. Step 2 Same choices

    Every model chooses from the same 70 assets.

  3. Step 3 One response, then locked

    No browsing, tools, or follow-up prompts. The model's submitted portfolio is frozen before results are known.

  4. Step 4 Fixed wait window

    The frozen portfolio sits untouched for 7 days or 1 month.

  5. Step 5 Prices score it

    Real ending prices decide which model did best.

Benchmark universe

What can models choose from?

The active roster, asset menu, horizons, and open rounds.

Models 10
Asset choices 70
Round lengths 2
Open rounds 23
Protocol Single-turn Non-agentic calls
Latest official results

What happened in the latest scored rounds?

Finished monthly and weekly rounds scored against real market returns.

Monthly result1 of 31
Monthly official result

Monthly result scored Aug 10

Same-window returns, ranked after final prices.

Scored
Model portfolios S&P 500 benchmark Maximum possible return
Claude Opus 4.8
Claude Opus 4.7
GPT-5.5
GPT-5.6 Sol
Grok 4.3
Claude Fable 5
Grok 4.5
Gemini 3.1 Pro
S&P 500
USO Crude Oil
Claude Opus 4.8 Anthropic
-0.42%
Claude Opus 4.7 Anthropic
-0.72%
GPT-5.5 OpenAI
-0.74%
GPT-5.6 Sol OpenAI
-0.95%
Grok 4.3 xAI
-1.42%
Claude Fable 5 Anthropic
-2.29%
Grok 4.5 xAI
-3.71%
Gemini 3.1 Pro Google
-7.19%
S&P 500 Benchmark
2.39%
USO Crude Oil - Hindsight best asset
15.84%
Portfolio context

Shows each model's saved portfolio weights.

Model portfolios

Ranked in the same order as the chart.

1
Claude Opus 4.8 Anthropic
Semiconductors (SMH) 30% Financials (XLF) 25% Industrials (XLI) 20% Taiwan (EWT) 15% Cybersecurity (CIBR) 10%
2
Claude Opus 4.7 Anthropic
Semiconductors (SMH) 30% Financials (XLF) 20% Biotechnology (XBI) 15% Cybersecurity (CIBR) 15% Taiwan (EWT) 20%
3
GPT-5.5 OpenAI
Semiconductors (SMH) 35% Taiwan (EWT) 20% Cybersecurity (CIBR) 20% Financials (XLF) 15% Biotechnology (XBI) 10%
4
GPT-5.6 Sol OpenAI
Semiconductors (SMH) 40% Cybersecurity (CIBR) 20% Biotechnology (XBI) 20% Financials (XLF) 10% Taiwan (EWT) 10%
5
Grok 4.3 xAI
Semiconductors (SMH) 50% Technology (XLK) 30% Cybersecurity (CIBR) 20%
6
Claude Fable 5 Anthropic
Semiconductors (SMH) 35% Taiwan (EWT) 15% Nasdaq 100 (QQQ) 20% Financials (XLF) 15% Industrials (XLI) 15%
7
Grok 4.5 xAI
Semiconductors (SMH) 40% Taiwan (EWT) 25% Technology (XLK) 20% Nasdaq 100 (QQQ) 15%
8
Gemini 3.1 Pro Google
Semiconductors (SMH) 40% South Korea (EWY) 30% Taiwan (EWT) 30%
Reference points

Not model portfolios.

S&P 500 Benchmark

Benchmark return over the same scoring window

USO Crude Oil - Hindsight best asset

100% Crude Oil (USO) hindsight ceiling

Official scored round

Monthly result scored Aug 10

Audit ID: CB-2026-07-10-1M

ScoredAug 10WindowJul 10 to Aug 10Models8Asset choices70LeaderClaude Opus 4.8HorizonMonthly
Live dashboard

What is still in progress?

Open portfolios, interim returns, and upcoming score dates.

Open rounds 23 19 monthly / 4 weekly
Frozen portfolios 181 37 assets currently held
Latest close Aug 12 Live returns update before final scoring
Next score Aug 13 Official results publish after ending prices
Audit packet

How can you verify the benchmark?

Round packets expose the report, prompt, portfolios, prices, hashes, and result status.