Weekly benchmark

Latest Result And Full Weekly History

Claude Fable 5 led the latest weekly round at 2.87%, compared with 0.47% for the S&P 500. Grok 4.3 leads the full weekly history with a CapitalBench Score of 4.7.

2026-08-21 close to 2026-08-28 close
Top model Claude Fable 5
Cumulative leader Grok 4.3
Completed weekly rounds 55
Top model return 2.87%
S&P 500 0.47%
Portfolio Minus S&P 500 +2.40 pp
CapitalBench Score 48.5 vs max possible · portfolio 2.87% · best asset 5.93%
CapitalBench Score

Main Benchmark Scale

CapitalBench Score compares each model portfolio with the best eligible asset in hindsight for the same weekly window. Max possible is always 100.

Score calculation
Latest round

Latest Round CapitalBench Score

CapitalBench Score compares each portfolio with the best eligible asset in hindsight for this exact weekly window. Max possible is 100.

Max possible S&P 500 Model portfolios
Max possible IGV · score 100
100.0 5.93%
S&P 500 Benchmark reference
8.0 0.47%
Claude Fable 5 Anthropic
48.5 2.87%
Claude Opus 5 Anthropic
41.8 2.48%
GPT-5.6 Sol OpenAI
41.8 2.48%
Grok 4.6 xAI
41.8 2.48%
Grok 4.3 xAI
9.6 0.57%
Gemini 3.1 Pro Google
5.8 0.34%
Grok 4.5 xAI
-1.3 -0.08%
Weekly track context

Cumulative Weekly Results

Models have beaten the S&P 500 in 34/55 completed weekly rounds and 173/356 model-round observations. At least one model beat the S&P 500 in the latest completed weekly round.

Open full weekly history
Completed rounds 55
Full-history leader Grok 4.3 4.7 score
S&P 500 score 2.5
Avg model return 0.09%
Avg vs S&P -0.17 pp
Model beat observations 173/356
Weekly history trend

CapitalBench Score By Completed Weekly Round

Max possible is fixed at 100. The chart compares S&P 500, the top model, and the average model score in each completed weekly round.

Max possible Top model Average model S&P 500
-250.0 0.0 50.0 100.0 CB-2026-05-24-1W: top model GPT-5.5 scored 38.9 CB-2026-05-24-1W: average model score 28.1 CB-2026-05-24-1W: S&P 500 score 11.1 May 29 05-24-1W CB-2026-05-27-1W: top model GPT-5.5 scored 48.4 CB-2026-05-27-1W: average model score 45.6 CB-2026-05-27-1W: S&P 500 score 10.5 Jun 2 05-27-1W CB-2026-05-28-1W: top model Gemini 3.1 Pro scored 85.1 CB-2026-05-28-1W: average model score 67.1 CB-2026-05-28-1W: S&P 500 score 7.1 Jun 4 05-28-1W CB-2026-05-29-1W: top model Claude Opus 4.8 scored -150.0 CB-2026-05-29-1W: average model score -176.1 CB-2026-05-29-1W: S&P 500 score -82.2 Jun 5 05-29-1W CB-2026-06-01-1W: top model Claude Opus 4.8 scored -143.8 CB-2026-06-01-1W: average model score -195.3 CB-2026-06-01-1W: S&P 500 score -82.2 Jun 5 06-01-1W CB-2026-06-02-1W: top model Claude Opus 4.8 scored -100.5 CB-2026-06-02-1W: average model score -210.1 CB-2026-06-02-1W: S&P 500 score -78.3 Jun 8 06-02-1W CB-2026-06-03-1W: top model Claude Opus 4.8 scored -128.6 CB-2026-06-03-1W: average model score -155.7 CB-2026-06-03-1W: S&P 500 score -53.1 Jun 9 06-03-1W CB-2026-06-05-1W: top model Claude Opus 4.7 scored 7.1 CB-2026-06-05-1W: average model score 1.1 CB-2026-06-05-1W: S&P 500 score 4.5 Jun 12 06-05-1W CB-2026-06-08-1W: top model Claude Opus 4.7 scored 12.7 CB-2026-06-08-1W: average model score -0.8 CB-2026-06-08-1W: S&P 500 score 15.2 Jun 15 06-08-1W CB-2026-06-09-1W: top model Grok 4.3 scored 6.6 CB-2026-06-09-1W: average model score 2.4 CB-2026-06-09-1W: S&P 500 score 15.2 Jun 16 06-09-1W CB-2026-06-12-1W: top model Claude Fable 5 scored 25.5 CB-2026-06-12-1W: average model score 3.5 CB-2026-06-12-1W: S&P 500 score 12.0 Jun 18 06-12-1W CB-2026-06-13-1W: top model GPT-5.5 scored 61.1 CB-2026-06-13-1W: average model score 35.9 CB-2026-06-13-1W: S&P 500 score 6.1 Jun 18 06-13-1W CB-2026-06-15-1W: top model GPT-5.5 scored 58.1 CB-2026-06-15-1W: average model score 38.2 CB-2026-06-15-1W: S&P 500 score -16.1 Jun 22 06-15-1W CB-2026-06-16-1W: top model Grok 4.3 scored -2.5 CB-2026-06-16-1W: average model score -18.9 CB-2026-06-16-1W: S&P 500 score -22.7 Jun 23 06-16-1W CB-2026-06-17-1W: top model Grok 4.3 scored 7.9 CB-2026-06-17-1W: average model score -13.0 CB-2026-06-17-1W: S&P 500 score -10.5 Jun 24 06-17-1W CB-2026-06-18-1W: top model Claude Opus 4.8 scored -27.7 CB-2026-06-18-1W: average model score -34.9 CB-2026-06-18-1W: S&P 500 score -21.3 Jun 25 06-18-1W CB-2026-06-22-1W: top model Claude Opus 4.8 scored -17.2 CB-2026-06-22-1W: average model score -37.5 CB-2026-06-22-1W: S&P 500 score -5.3 Jun 29 06-22-1W CB-2026-06-23-1W: top model GPT-5.5 scored 58.6 CB-2026-06-23-1W: average model score 52.4 CB-2026-06-23-1W: S&P 500 score 23.6 Jun 30 06-23-1W CB-2026-06-24-1W: top model Claude Opus 4.8 scored 25.8 CB-2026-06-24-1W: average model score 19.2 CB-2026-06-24-1W: S&P 500 score 19.4 Jul 1 06-24-1W CB-2026-06-25-1W: top model Grok 4.3 scored 36.2 CB-2026-06-25-1W: average model score 11.9 CB-2026-06-25-1W: S&P 500 score 13.7 Jul 2 06-25-1W CB-2026-06-26-1W: top model Grok 4.3 scored 23.0 CB-2026-06-26-1W: average model score 16.9 CB-2026-06-26-1W: S&P 500 score 26.6 Jul 2 06-26-1W CB-2026-06-29-1W: top model Claude Opus 4.8 scored 2.5 CB-2026-06-29-1W: average model score -8.6 CB-2026-06-29-1W: S&P 500 score 13.0 Jul 6 06-29-1W CB-2026-06-30-1W: top model Claude Opus 4.7 scored -20.3 CB-2026-06-30-1W: average model score -29.5 CB-2026-06-30-1W: S&P 500 score 0.9 Jul 7 06-30-1W CB-2026-07-01-1W: top model Grok 4.3 scored 25.0 CB-2026-07-01-1W: average model score 9.7 CB-2026-07-01-1W: S&P 500 score -0.6 Jul 8 07-01-1W CB-2026-07-02-1W: top model Gemini 3.1 Pro scored 3.2 CB-2026-07-02-1W: average model score -10.8 CB-2026-07-02-1W: S&P 500 score 19.2 Jul 9 07-02-1W CB-2026-07-06-1W: top model Grok 4.3 scored -12.7 CB-2026-07-06-1W: average model score -20.2 CB-2026-07-06-1W: S&P 500 score -2.2 Jul 13 07-06-1W CB-2026-07-07-1W: top model GPT-5.5 scored 41.0 CB-2026-07-07-1W: average model score 9.0 CB-2026-07-07-1W: S&P 500 score 5.3 Jul 14 07-07-1W CB-2026-07-08-1W: top model Gemini 3.1 Pro scored 43.0 CB-2026-07-08-1W: average model score 13.4 CB-2026-07-08-1W: S&P 500 score 11.7 Jul 15 07-08-1W CB-2026-07-09-1W: top model Claude Opus 4.8 scored -22.6 CB-2026-07-09-1W: average model score -39.4 CB-2026-07-09-1W: S&P 500 score -1.4 Jul 16 07-09-1W CB-2026-07-10-1W: top model Claude Opus 4.7 scored -32.0 CB-2026-07-10-1W: average model score -42.2 CB-2026-07-10-1W: S&P 500 score -11.0 Jul 17 07-10-1W CB-2026-07-13-1W: top model Grok 4.3 scored 91.3 CB-2026-07-13-1W: average model score 43.4 CB-2026-07-13-1W: S&P 500 score -13.2 Jul 20 07-13-1W CB-2026-07-14-1W: top model Grok 4.3 scored 49.6 CB-2026-07-14-1W: average model score 23.5 CB-2026-07-14-1W: S&P 500 score -6.5 Jul 21 07-14-1W CB-2026-07-15-1W: top model Grok 4.3 scored 78.2 CB-2026-07-15-1W: average model score 43.8 CB-2026-07-15-1W: S&P 500 score -11.6 Jul 22 07-15-1W CB-2026-07-17-1W: top model Gemini 3.1 Pro scored 27.1 CB-2026-07-17-1W: average model score 10.9 CB-2026-07-17-1W: S&P 500 score -5.7 Jul 24 07-17-1W CB-2026-07-20-1W: top model Claude Fable 5 scored 13.1 CB-2026-07-20-1W: average model score 4.6 CB-2026-07-20-1W: S&P 500 score -6.3 Jul 27 07-20-1W CB-2026-07-21-1W: top model GPT-5.6 Sol scored 13.9 CB-2026-07-21-1W: average model score -18.5 CB-2026-07-21-1W: S&P 500 score -14.9 Jul 28 07-21-1W CB-2026-07-22-1W: top model GPT-5.6 Sol scored 6.9 CB-2026-07-22-1W: average model score -7.7 CB-2026-07-22-1W: S&P 500 score -56.2 Jul 29 07-22-1W CB-2026-07-23-1W: top model Grok 4.3 scored 5.4 CB-2026-07-23-1W: average model score -5.5 CB-2026-07-23-1W: S&P 500 score 6.7 Jul 30 07-23-1W CB-2026-07-24-1W: top model GPT-5.6 Sol scored 23.9 CB-2026-07-24-1W: average model score 4.3 CB-2026-07-24-1W: S&P 500 score 14.6 Jul 31 07-24-1W CB-2026-07-27-1W: top model Grok 4.3 scored 19.5 CB-2026-07-27-1W: average model score 9.4 CB-2026-07-27-1W: S&P 500 score 35.1 Aug 3 07-27-1W CB-2026-07-28-1W: top model Grok 4.3 scored 11.5 CB-2026-07-28-1W: average model score 0.2 CB-2026-07-28-1W: S&P 500 score 31.6 Aug 4 07-28-1W CB-2026-07-29-1W: top model Claude Opus 5 scored 9.1 CB-2026-07-29-1W: average model score 3.2 CB-2026-07-29-1W: S&P 500 score 32.0 Aug 5 07-29-1W CB-2026-07-30-1W: top model Gemini 3.1 Pro scored 41.6 CB-2026-07-30-1W: average model score 23.5 CB-2026-07-30-1W: S&P 500 score 42.7 Aug 6 07-30-1W CB-2026-07-31-1W: top model Grok 4.3 scored 23.4 CB-2026-07-31-1W: average model score 7.8 CB-2026-07-31-1W: S&P 500 score 23.4 Aug 7 07-31-1W CB-2026-08-04-1W: top model GPT-5.5 scored 28.6 CB-2026-08-04-1W: average model score 7.2 CB-2026-08-04-1W: S&P 500 score -1.0 Aug 11 08-04-1W CB-2026-08-05-1W: top model Grok 4.5 scored 27.1 CB-2026-08-05-1W: average model score 12.5 CB-2026-08-05-1W: S&P 500 score 3.2 Aug 12 08-05-1W CB-2026-08-07-1W: top model GPT-5.5 scored 48.9 CB-2026-08-07-1W: average model score 24.3 CB-2026-08-07-1W: S&P 500 score 4.8 Aug 14 08-07-1W CB-2026-08-09-1W: top model GPT-5.6 Sol scored 10.2 CB-2026-08-09-1W: average model score 4.5 CB-2026-08-09-1W: S&P 500 score -0.3 Aug 17 08-09-1W CB-2026-08-11-1W: top model GPT-5.6 Sol scored 61.3 CB-2026-08-11-1W: average model score 30.3 CB-2026-08-11-1W: S&P 500 score -8.9 Aug 18 08-11-1W CB-2026-08-13-1W: top model Gemini 3.1 Pro scored 14.5 CB-2026-08-13-1W: average model score 7.1 CB-2026-08-13-1W: S&P 500 score -8.4 Aug 20 08-13-1W CB-2026-08-15-1W: top model Claude Opus 5 scored 10.2 CB-2026-08-15-1W: average model score 5.0 CB-2026-08-15-1W: S&P 500 score -4.0 Aug 24 08-15-1W CB-2026-08-18-1W: top model Grok 4.3 scored 10.2 CB-2026-08-18-1W: average model score -0.0 CB-2026-08-18-1W: S&P 500 score -0.7 Aug 25 08-18-1W CB-2026-08-19-1W: top model Claude Opus 5 scored 13.5 CB-2026-08-19-1W: average model score 4.7 CB-2026-08-19-1W: S&P 500 score -2.2 Aug 26 08-19-1W CB-2026-08-20-1W: top model Claude Fable 5 scored 52.6 CB-2026-08-20-1W: average model score 33.7 CB-2026-08-20-1W: S&P 500 score 11.2 Aug 27 08-20-1W CB-2026-08-21-1W: top model Claude Fable 5 scored 48.5 CB-2026-08-21-1W: average model score 26.9 CB-2026-08-21-1W: S&P 500 score 8.0 Aug 28 08-21-1W
Investor readout

What This Weekly Round Shows

This section summarizes the benchmark-relative result, model crowding, and the hindsight asset that set the maximum possible return for the round.

Benchmark result Claude Fable 5 +2.40 pp

Top model return 2.87% vs S&P 500 0.47%.

Model average 1.59%

Average model portfolio finished +1.12 pp versus S&P 500.

Most crowded model exposures
Cybersecurity (CIBR) 30% avg Aerospace and Defense (ITA) 28.6% avg Software (IGV) 17.9% avg Ethereum ETF (ETHA) 5% avg
What worked in hindsight
Software (IGV) 5.93% Cybersecurity (CIBR) 3.91% Taiwan Equities (EWT) 3.45%
What hurt the model portfolios
Aerospace and Defense (ITA) -1.90% Autonomous Technology and Robotics (ARKQ) -2.60% Semiconductors (SMH) -1.30%
Official weekly result history

Browse Weekly Official Results

Move backward or forward through completed weekly rounds without leaving the weekly results page.

Weekly result1 of 55
Weekly official result

Weekly result scored Aug 28

Same-window returns, ranked after final prices.

Scored
Model portfolios S&P 500 benchmark Maximum possible return
Claude Fable 5
GPT-5.6 Sol
Claude Opus 5
Grok 4.6
Grok 4.3
Gemini 3.1 Pro
Grok 4.5
S&P 500
IGV Software
Claude Fable 5 Anthropic
2.87%
GPT-5.6 Sol OpenAI
2.48%
Claude Opus 5 Anthropic
2.48%
Grok 4.6 xAI
2.48%
Grok 4.3 xAI
0.57%
Gemini 3.1 Pro Google
0.34%
Grok 4.5 xAI
-0.08%
S&P 500 Benchmark
0.47%
IGV Software - Hindsight best asset
5.93%
Portfolio context

Shows each model's saved portfolio weights.

Model portfolios

Ranked in the same order as the chart.

1
Claude Fable 5 Anthropic
Cybersecurity (CIBR) 35% Software (IGV) 35% Defense (ITA) 30%
2
GPT-5.6 Sol OpenAI
Cybersecurity (CIBR) 35% Defense (ITA) 35% Software (IGV) 30%
3
Claude Opus 5 Anthropic
Cybersecurity (CIBR) 35% Defense (ITA) 35% Software (IGV) 30%
4
Grok 4.6 xAI
Cybersecurity (CIBR) 35% Defense (ITA) 35% Software (IGV) 30%
5
Grok 4.3 xAI
Ethereum ETF (ETHA) 35% Bitcoin ETF (IBIT) 35% S&P 500 (SPY) 30%
6
Gemini 3.1 Pro Google
Semiconductors (SMH) 35% Cybersecurity (CIBR) 35% Defense (ITA) 30%
7
Grok 4.5 xAI
Cybersecurity (CIBR) 35% Defense (ITA) 35% Robotics (ARKQ) 30%
Reference points

Not model portfolios.

S&P 500 Benchmark

Benchmark return over the same scoring window

IGV Software - Hindsight best asset

100% Software (IGV) hindsight ceiling

Official scored round

Weekly result scored Aug 28

Audit ID: CB-2026-08-21-1W

ScoredAug 28WindowAug 21 to Aug 28Models7Asset choices70LeaderClaude Fable 5HorizonWeekly
Result insights

What The Latest Weekly Result Shows

Context generated from the scored weekly round, including consensus performance, oracle comparison, and benchmark difficulty.

Consensus PerformanceAug 21-Aug 28

AI consensus portfolio scored 26.9 versus the oracle

If the weekly model allocations were averaged into one consensus portfolio, it returned +1.59% versus +0.47% for the S&P 500 and +5.93% for the hindsight best asset.

Consensus means the average of model allocations in the same round. CapitalBench Score compares that return with the hindsight-best eligible asset for that exact scoring window.

High confidenceMath: deterministicData through Aug 28, 2026
Consensus Portfolio Return
+1.59%
Average Model Return
+1.59%
Consensus Capitalbench Score
26.9
Why it matters

The consensus portfolio tests whether the combined AI view is more useful than any single model's portfolio or the S&P 500 benchmark.

Benchmark DifficultyAug 21-Aug 28

Weekly round had +10.23% asset dispersion

The best scored asset returned +5.93%, the worst returned -4.30%, and +40.00% of the universe was positive. The S&P 500 ranked 17 out of 70 options.

Asset dispersion is the gap between the best and worst eligible assets in the same round. Wider dispersion makes missed allocation choices more costly.

High confidenceMath: deterministicData through Aug 28, 2026
Oracle Return
+5.93%
Worst Asset Return
-4.30%
Positive Universe Share
+40.0%
Why it matters

Benchmark difficulty matters because model scores should be interpreted against the opportunity set and the market window they faced.

Oracle ComparisonAug 21-Aug 28

Models found the weekly oracle asset

The hindsight best asset was Software (IGV) at +5.93%. 4 of 7 models held it, with +17.86% average allocation. The largest allocation came from Claude Fable 5 at +35.00%.

Oracle means the best eligible asset in hindsight for that round. Models do not know it when portfolios are frozen.

High confidenceMath: deterministicData through Aug 28, 2026
Oracle Asset Holder Count
4
Average Oracle Asset Allocation
+17.9%
Why it matters

This shows whether models identified the eventual best asset before scoring, even when portfolio weights were too small to fully capture the oracle return.

Detailed scores

Weekly Score Tables

Switch between the latest completed weekly round and cumulative weekly history. All cumulative rows are derived from resolved weekly run files.

CapitalBench Score audit

Score Formula

100 matches the maximum possible return; negative scores preserve the size of a loss.

Claude Fable 5 score 48.5 · portfolio 2.87% · max possible 5.93%: Software (IGV) GPT-5.6 Sol score 41.8 · portfolio 2.48% · max possible 5.93%: Software (IGV) Claude Opus 5 score 41.8 · portfolio 2.48% · max possible 5.93%: Software (IGV) Grok 4.6 score 41.8 · portfolio 2.48% · max possible 5.93%: Software (IGV) Grok 4.3 score 9.6 · portfolio 0.57% · max possible 5.93%: Software (IGV) Gemini 3.1 Pro score 5.8 · portfolio 0.34% · max possible 5.93%: Software (IGV) Grok 4.5 score -1.3 · portfolio -0.08% · max possible 5.93%: Software (IGV)
Claude Fable 5AnthropicCYBERSECURITY32.87%0.47%+2.40 pp3.06%
GPT-5.6 SolOpenAICYBERSECURITY32.48%0.47%+2.01 pp3.45%
Claude Opus 5AnthropicCYBERSECURITY32.48%0.47%+2.01 pp3.45%
Grok 4.6xAICYBERSECURITY32.48%0.47%+2.01 pp3.45%
Grok 4.3xAIETHEREUM_ETF30.57%0.47%+0.09 pp5.36%
Gemini 3.1 ProGoogleSEMICONDUCTORS30.34%0.47%-0.13 pp5.59%
Grok 4.5xAICYBERSECURITY3-0.08%0.47%-0.55 pp6.01%
Portfolios behind the result

Portfolios Behind The Result

Compact allocation cards first, with the full searchable table below.

Anthropic
Claude Fable 5 Anthropic
Cybersecurity (CIBR) 35% Software (IGV) 35% Aerospace and Defense (ITA) 30%
OpenAI
GPT-5.6 Sol OpenAI
Cybersecurity (CIBR) 35% Aerospace and Defense (ITA) 35% Software (IGV) 30%
Anthropic
Claude Opus 5 Anthropic
Cybersecurity (CIBR) 35% Aerospace and Defense (ITA) 35% Software (IGV) 30%
xAI
Grok 4.6 xAI
Cybersecurity (CIBR) 35% Aerospace and Defense (ITA) 35% Software (IGV) 30%
xAI
Grok 4.3 xAI
Ethereum ETF (ETHA) 35% Bitcoin ETF (IBIT) 35% S&P 500 (SPY) 30%
Google
Gemini 3.1 Pro Google
Semiconductors (SMH) 35% Cybersecurity (CIBR) 35% Aerospace and Defense (ITA) 30%
xAI
Grok 4.5 xAI
Cybersecurity (CIBR) 35% Aerospace and Defense (ITA) 35% Autonomous Technology and Robotics (ARKQ) 30%
ModelProviderPortfolioConfidenceProtocol
Claude Fable 5
anthropic-claude-fable-5
Anthropic
Cybersecurity (CIBR)35%Software (IGV)35%Aerospace and Defense (ITA)30%
0.58
Portfolio round
GPT-5.6 Sol
openai-gpt-5-6-sol
OpenAI
Cybersecurity (CIBR)35%Aerospace and Defense (ITA)35%Software (IGV)30%
0.58
Portfolio round
Claude Opus 5
anthropic-claude-opus-5
Anthropic
Cybersecurity (CIBR)35%Aerospace and Defense (ITA)35%Software (IGV)30%
0.56
Portfolio round
Grok 4.6
xai-grok-4-6
xAI
Cybersecurity (CIBR)35%Aerospace and Defense (ITA)35%Software (IGV)30%
0.56
Portfolio round
Grok 4.3
xai-grok-4-3
xAI
Ethereum ETF (ETHA)35%Bitcoin ETF (IBIT)35%S&P 500 (SPY)30%
0.60
Portfolio round
Gemini 3.1 Pro
google-gemini-3-1-pro
Google
Semiconductors (SMH)35%Cybersecurity (CIBR)35%Aerospace and Defense (ITA)30%
0.59
Portfolio round
Grok 4.5
xai-grok-4-5
xAI
Cybersecurity (CIBR)35%Aerospace and Defense (ITA)35%Autonomous Technology and Robotics (ARKQ)30%
0.59
Portfolio round
Current live weekly round CB-2026-08-30-1W Scores after the 2026-09-08 close.
View current portfolios