Insights

Readable signals from the AI capital allocation benchmark

Daily findings from model portfolios, scoring windows, AI Risk Appetite, benchmark difficulty, consensus positioning, and model behavior.

Published insights 27

Readable signals in the latest generated feed.

Data through Jul 30, 2026

Most recent close or result date used by the engine.

High confidence 10

Findings backed by deterministic calculations and direct evidence.

Deterministic math 27

Insights generated without LLM interpretation.

Latest feed

What The Benchmark Is Showing Now

Each card includes the calculation source, evidence links, and why the signal may matter to investors, allocators, traders, and AI researchers.

API Docs
Consensus Performance Jul 23-Jul 30
Weekly result CB-2026-07-23-1W 7 models Oracle: Software (IGV), +7.13% Resolved result

AI consensus portfolio scored -5.5 versus the oracle

If the weekly model allocations were averaged into one consensus portfolio, it returned -0.39% versus +0.48% for the S&P 500 and +7.13% for the hindsight best asset.

Consensus means the average of model allocations in the same round. CapitalBench Score compares that return with the hindsight-best eligible asset for that exact scoring window.

High confidenceMath: deterministicData through Jul 30, 2026
Consensus Portfolio Return
-0.39%
Average Model Return
-0.39%
Consensus Capitalbench Score
-5.5
Consensus Performance Jun 30-Jul 30
Monthly result CB-2026-06-30-1M 5 models Oracle: Ethereum ETF (ETHA), +22.04% Resolved result

AI consensus portfolio scored -23.6 versus the oracle

If the monthly model allocations were averaged into one consensus portfolio, it returned -5.19% versus -0.68% for the S&P 500 and +22.04% for the hindsight best asset.

Consensus means the average of model allocations in the same round. CapitalBench Score compares that return with the hindsight-best eligible asset for that exact scoring window.

High confidenceMath: deterministicData through Jul 30, 2026
Consensus Portfolio Return
-5.19%
Average Model Return
-5.19%
Consensus Capitalbench Score
-23.6
Benchmark Difficulty Jul 23-Jul 30
Weekly result CB-2026-07-23-1W 7 models Oracle: Software (IGV), +7.13% Resolved result

Weekly round had +15.74% asset dispersion

The best scored asset returned +7.13%, the worst returned -8.61%, and +67.14% of the universe was positive. The S&P 500 ranked 32 out of 70 options.

Asset dispersion is the gap between the best and worst eligible assets in the same round. Wider dispersion makes missed allocation choices more costly.

High confidenceMath: deterministicData through Jul 30, 2026
Oracle Return
+7.13%
Worst Asset Return
-8.61%
Positive Universe Share
+67.1%
Benchmark Difficulty Jun 30-Jul 30
Monthly result CB-2026-06-30-1M 5 models Oracle: Ethereum ETF (ETHA), +22.04% Resolved result

Monthly round had +42.19% asset dispersion

The best scored asset returned +22.04%, the worst returned -20.15%, and +50.00% of the universe was positive. The S&P 500 ranked 41 out of 70 options.

Asset dispersion is the gap between the best and worst eligible assets in the same round. Wider dispersion makes missed allocation choices more costly.

High confidenceMath: deterministicData through Jul 30, 2026
Oracle Return
+22.0%
Worst Asset Return
-20.2%
Positive Universe Share
+50.0%
Oracle Comparison Jul 23-Jul 30
Weekly result CB-2026-07-23-1W 7 models Oracle: Software (IGV), +7.13% Resolved result

Models missed the weekly oracle asset

The hindsight best asset was Software (IGV) at +7.13%. 0 of 7 models held it, with +0.00% average allocation.

Oracle means the best eligible asset in hindsight for that round. Models do not know it when portfolios are frozen.

High confidenceMath: deterministicData through Jul 30, 2026
Oracle Asset Holder Count
0.00
Average Oracle Asset Allocation
0.00
Oracle Comparison Jun 30-Jul 30
Monthly result CB-2026-06-30-1M 5 models Oracle: Ethereum ETF (ETHA), +22.04% Resolved result

Models missed the monthly oracle asset

The hindsight best asset was Ethereum ETF (ETHA) at +22.04%. 0 of 5 models held it, with +0.00% average allocation.

Oracle means the best eligible asset in hindsight for that round. Models do not know it when portfolios are frozen.

High confidenceMath: deterministicData through Jul 30, 2026
Oracle Asset Holder Count
0.00
Average Oracle Asset Allocation
0.00
Confidence Calibration As of Jul 30
All resolved official results 62 resolved rounds 342 scored results Median confidence 0.58 Resolved history

High-confidence model calls have underperformed lower-confidence calls

Across resolved official results, submissions at or above the median confidence of 0.58 averaged -1.66%, while lower-confidence submissions averaged -0.58%.

Confidence is the model's own 0-1 self-reported confidence at submission time, compared with later realized returns.

High confidenceMath: deterministicData through Jul 30, 2026
High Confidence Average Return
-1.66%
Low Confidence Average Return
-0.58%
High Confidence Average Capitalbench Score
-16.3
Performance Attribution Jul 23-Jul 30
Weekly result CB-2026-07-23-1W 7 models Model: Grok 4.3 Resolved result

Grok 4.3's result was driven by Financials Sector

In the latest weekly result, Financials Sector contributed +0.42% to Grok 4.3's portfolio. The largest drag came from Energy Sector at -0.28%.

Attribution multiplies each frozen holding's weight by its asset return to show what helped or hurt the model portfolio.

High confidenceMath: deterministicData through Jul 30, 2026
Largest Positive Contribution
+0.42%
Largest Negative Contribution
-0.28%
Performance Attribution Jun 30-Jul 30
Monthly result CB-2026-06-30-1M 5 models Model: Claude Opus 4.8 Resolved result

Claude Opus 4.8's result was driven by Financials Sector

In the latest monthly result, Financials Sector contributed +1.26% to Claude Opus 4.8's portfolio. The largest drag came from Industrials Sector at -0.92%.

Attribution multiplies each frozen holding's weight by its asset return to show what helped or hurt the model portfolio.

High confidenceMath: deterministicData through Jul 30, 2026
Largest Positive Contribution
+1.26%
Largest Negative Contribution
-0.92%
Model Behavior As of Jul 30
Model behavior profiles 9 models

Model allocation styles are separating into clear behavior profiles

GPT-5.5 has the highest average risk-taking score at 80.0/100. Grok 4.3 has the largest average top holding at +40.23%. Claude Opus 5 has the lowest measured turnover at +40.62%.

Momentum exposure measures how much of the frozen portfolio went into assets that had already been recent winners before the model made its allocation.

High confidenceMath: deterministicData through Jul 30, 2026
Highest Average Risk Taking Score
80.0/100
Largest Average Top Holding
40.2
Lowest Average Turnover
40.6
Market Environment As of Jul 30
Weekly market environments 10 resolved rounds 2 models Ready sample

Weekly model leadership changes with the S&P 500 environment

Grok 4.3 leads down environments at +2.35% across 4 tests; GPT-5.5 leads up environments at +2.35% across 6 tests.

Market environments group resolved rounds by the S&P 500 return over the same weekly or monthly window. Models are compared only on shared rounds; high confidence requires at least six observations and stable leadership.

Medium confidenceMath: deterministicData through Jul 30, 2026
Down Leader Average Return
+2.35%
Down Shared Rounds
4
Down Leader Stability
1.00
Risk Regime As of Jul 30
Latest live portfolios 2 live rounds 16 models Live portfolios

Live AI risk posture is risk-seeking

The newest live portfolios have a deterministic risk-taking score of 69.1 out of 100.

Risk-taking score is allocation-based, not performance-based: higher means more weight in growth, momentum, cyclical, and higher-risk assets.

Medium confidenceMath: deterministicData through Jul 30, 2026
Live Risk Taking Score
69.1/100
Market Environment As of Jul 30
Monthly market environments 12 resolved rounds 2 models Ready sample

Monthly model leadership changes with the S&P 500 environment

Grok 4.3 leads down environments at -1.22% across 6 tests; Claude Opus 4.8 leads up environments at +1.34% across 6 tests.

Market environments group resolved rounds by the S&P 500 return over the same weekly or monthly window. Models are compared only on shared rounds; high confidence requires at least six observations and stable leadership.

Medium confidenceMath: deterministicData through Jul 30, 2026
Down Leader Average Return
-1.22%
Down Shared Rounds
6
Down Leader Stability
1.00
Horizon Agreement As of Jul 30
Latest live portfolios 2 live rounds 16 models Live portfolios

Weekly and monthly AI portfolios both favor broad and cyclical equity

The newest weekly portfolios allocate +56.88% to broad and cyclical equity, while the newest monthly portfolios allocate +65.62%.

Horizon agreement compares the newest weekly and monthly live portfolios to see whether short- and longer-window model stances line up.

Medium confidenceMath: deterministicData through Jul 30, 2026
Weekly Top Regime Allocation
56.9
Monthly Top Regime Allocation
65.6
Current Positioning As of Jul 30
Latest live portfolios 2 live rounds 16 models Live portfolios

Live AI portfolios are concentrated in Financials Sector (XLF)

Across the newest live weekly and monthly portfolios, Financials Sector (XLF) is the largest aggregate allocation at +21.88%.

Aggregate allocation averages the newest live model portfolios before final scores are known.

Medium confidenceMath: deterministicData through Jul 30, 2026
Aggregate Live Allocation
21.9
Live Performance As of Jul 30
Open-round interim performance 22 open rounds 9 models Interim, not final

GPT-5.5 has the strongest live alpha

Using the latest available interim close, GPT-5.5 in CB-2026-07-13-1M is ahead of the S&P 500 by +6.15 percentage points, while Gemini 3.1 Pro in CB-2026-07-10-1M is at -10.06 percentage points.

Live alpha is interim model return minus interim S&P 500 return. It is provisional until the round reaches its official score date.

Medium confidenceMath: deterministicData through Jul 30, 2026
Best Live Alpha
6.15
Worst Live Alpha
-10.1
Model Similarity As of Jul 30
Latest live portfolios 2 live rounds 16 models Live portfolios

Live model portfolios are tightly clustered

The closest live allocation pair is Claude Opus 5 and Gemini 3.1 Pro with +91.39% cosine similarity. The current allocation outlier is GPT-5.6 Sol.

Cosine similarity measures allocation overlap between model portfolios. A value near 1.00 means the weights are very similar.

Medium confidenceMath: deterministicData through Jul 30, 2026
Closest Pair Cosine Similarity
0.91
Outlier Average Distance
0.69
Insight families

What The Engine Looks For

The engine is designed to surface useful behavior and performance patterns, not generic market commentary.

12 signals

Market Environment

Which models lead, remain consistent, or change most across resolved down, flat, and up S&P 500 environments.

2 signals

Benchmark Difficulty

How hard a scoring window was, based on the spread between the best, worst, and broad market outcomes.

2 signals

Consensus Performance

Whether the average AI portfolio performed well against the S&P 500 and the hindsight-best asset in the same round.

2 signals

Oracle Comparison

Whether models found, missed, or underweighted the asset that later turned out to be best.

2 signals

Performance Attribution

Which holdings drove a model's realized result after the frozen portfolio was scored.

1 signal

Confidence Calibration

Whether self-reported model confidence has contained useful information about later results.

Method

How Insights Are Produced

Deterministic calculations are the source of truth. LLM-assisted wording can polish selected titles and summaries, but it cannot change calculations, evidence links, round context, or benchmark facts.

Latest generation: Jul 31, 2026, 5:33 PM UTC

  1. 1 Build the input packet

    Collect public rounds, official portfolios, results, live marks, asset risk ratings, and benchmark sets.

  2. 2 Run deterministic math

    Calculate consensus performance, benchmark difficulty, market-environment results, risk posture, similarity, attribution, and live paths.

  3. 3 Attach evidence

    Every insight links back to round pages, leaderboard pages, scoring files, or methodology pages.

  4. 4 Validate before publishing

    The feed must pass schema checks before the website and API expose it.