Readable signals in the latest generated feed.
Readable signals from the AI capital allocation benchmark
Daily findings from model portfolios, scoring windows, AI Risk Appetite, benchmark difficulty, consensus positioning, and model behavior.
Most recent close or result date used by the engine.
Findings backed by deterministic calculations and direct evidence.
Insights generated without LLM interpretation.
What The Benchmark Is Showing Now
Each card includes the calculation source, evidence links, and why the signal may matter to investors, allocators, traders, and AI researchers.
AI consensus portfolio scored -5.5 versus the oracle
If the weekly model allocations were averaged into one consensus portfolio, it returned -0.39% versus +0.48% for the S&P 500 and +7.13% for the hindsight best asset.
Consensus means the average of model allocations in the same round. CapitalBench Score compares that return with the hindsight-best eligible asset for that exact scoring window.
- Consensus Portfolio Return
- -0.39%
- Average Model Return
- -0.39%
- Consensus Capitalbench Score
- -5.5
AI consensus portfolio scored -23.6 versus the oracle
If the monthly model allocations were averaged into one consensus portfolio, it returned -5.19% versus -0.68% for the S&P 500 and +22.04% for the hindsight best asset.
Consensus means the average of model allocations in the same round. CapitalBench Score compares that return with the hindsight-best eligible asset for that exact scoring window.
- Consensus Portfolio Return
- -5.19%
- Average Model Return
- -5.19%
- Consensus Capitalbench Score
- -23.6
Weekly round had +15.74% asset dispersion
The best scored asset returned +7.13%, the worst returned -8.61%, and +67.14% of the universe was positive. The S&P 500 ranked 32 out of 70 options.
Asset dispersion is the gap between the best and worst eligible assets in the same round. Wider dispersion makes missed allocation choices more costly.
- Oracle Return
- +7.13%
- Worst Asset Return
- -8.61%
- Positive Universe Share
- +67.1%
Monthly round had +42.19% asset dispersion
The best scored asset returned +22.04%, the worst returned -20.15%, and +50.00% of the universe was positive. The S&P 500 ranked 41 out of 70 options.
Asset dispersion is the gap between the best and worst eligible assets in the same round. Wider dispersion makes missed allocation choices more costly.
- Oracle Return
- +22.0%
- Worst Asset Return
- -20.2%
- Positive Universe Share
- +50.0%
Models missed the weekly oracle asset
The hindsight best asset was Software (IGV) at +7.13%. 0 of 7 models held it, with +0.00% average allocation.
Oracle means the best eligible asset in hindsight for that round. Models do not know it when portfolios are frozen.
- Oracle Asset Holder Count
- 0.00
- Average Oracle Asset Allocation
- 0.00
Models missed the monthly oracle asset
The hindsight best asset was Ethereum ETF (ETHA) at +22.04%. 0 of 5 models held it, with +0.00% average allocation.
Oracle means the best eligible asset in hindsight for that round. Models do not know it when portfolios are frozen.
- Oracle Asset Holder Count
- 0.00
- Average Oracle Asset Allocation
- 0.00
High-confidence model calls have underperformed lower-confidence calls
Across resolved official results, submissions at or above the median confidence of 0.58 averaged -1.66%, while lower-confidence submissions averaged -0.58%.
Confidence is the model's own 0-1 self-reported confidence at submission time, compared with later realized returns.
- High Confidence Average Return
- -1.66%
- Low Confidence Average Return
- -0.58%
- High Confidence Average Capitalbench Score
- -16.3
Grok 4.3's result was driven by Financials Sector
In the latest weekly result, Financials Sector contributed +0.42% to Grok 4.3's portfolio. The largest drag came from Energy Sector at -0.28%.
Attribution multiplies each frozen holding's weight by its asset return to show what helped or hurt the model portfolio.
- Largest Positive Contribution
- +0.42%
- Largest Negative Contribution
- -0.28%
Claude Opus 4.8's result was driven by Financials Sector
In the latest monthly result, Financials Sector contributed +1.26% to Claude Opus 4.8's portfolio. The largest drag came from Industrials Sector at -0.92%.
Attribution multiplies each frozen holding's weight by its asset return to show what helped or hurt the model portfolio.
- Largest Positive Contribution
- +1.26%
- Largest Negative Contribution
- -0.92%
Model allocation styles are separating into clear behavior profiles
GPT-5.5 has the highest average risk-taking score at 80.0/100. Grok 4.3 has the largest average top holding at +40.23%. Claude Opus 5 has the lowest measured turnover at +40.62%.
Momentum exposure measures how much of the frozen portfolio went into assets that had already been recent winners before the model made its allocation.
- Highest Average Risk Taking Score
- 80.0/100
- Largest Average Top Holding
- 40.2
- Lowest Average Turnover
- 40.6
Weekly model leadership changes with the S&P 500 environment
Grok 4.3 leads down environments at +2.35% across 4 tests; GPT-5.5 leads up environments at +2.35% across 6 tests.
Market environments group resolved rounds by the S&P 500 return over the same weekly or monthly window. Models are compared only on shared rounds; high confidence requires at least six observations and stable leadership.
- Down Leader Average Return
- +2.35%
- Down Shared Rounds
- 4
- Down Leader Stability
- 1.00
Live AI risk posture is risk-seeking
The newest live portfolios have a deterministic risk-taking score of 69.1 out of 100.
Risk-taking score is allocation-based, not performance-based: higher means more weight in growth, momentum, cyclical, and higher-risk assets.
- Live Risk Taking Score
- 69.1/100
Monthly model leadership changes with the S&P 500 environment
Grok 4.3 leads down environments at -1.22% across 6 tests; Claude Opus 4.8 leads up environments at +1.34% across 6 tests.
Market environments group resolved rounds by the S&P 500 return over the same weekly or monthly window. Models are compared only on shared rounds; high confidence requires at least six observations and stable leadership.
- Down Leader Average Return
- -1.22%
- Down Shared Rounds
- 6
- Down Leader Stability
- 1.00
Weekly and monthly AI portfolios both favor broad and cyclical equity
The newest weekly portfolios allocate +56.88% to broad and cyclical equity, while the newest monthly portfolios allocate +65.62%.
Horizon agreement compares the newest weekly and monthly live portfolios to see whether short- and longer-window model stances line up.
- Weekly Top Regime Allocation
- 56.9
- Monthly Top Regime Allocation
- 65.6
Grok 4.3 leads when the S&P 500 is negative
The model averaged +2.35% across 4 shared-cohort rounds drawn from 14 resolved weekly down environments. The sample meets publication thresholds.
Market environments group resolved rounds by the S&P 500 return over the same weekly or monthly window. Models are compared only on shared rounds; high confidence requires at least six observations and stable leadership.
- Average Model Return
- +2.35%
- Average S&P 500 Return
- -1.01%
- Capitalbench Score
- 23.5
GPT-5.5 leads when the S&P 500 is positive
The model averaged +2.35% across 6 shared-cohort rounds drawn from 15 resolved weekly up environments. The sample meets publication thresholds.
Market environments group resolved rounds by the S&P 500 return over the same weekly or monthly window. Models are compared only on shared rounds; high confidence requires at least six observations and stable leadership.
- Average Model Return
- +2.35%
- Average S&P 500 Return
- +1.07%
- Capitalbench Score
- 23.9
Grok 4.3 leads when the S&P 500 is negative
The model averaged -1.22% across 6 shared-cohort rounds drawn from 8 resolved monthly down environments. The sample meets publication thresholds.
Market environments group resolved rounds by the S&P 500 return over the same weekly or monthly window. Models are compared only on shared rounds; high confidence requires at least six observations and stable leadership.
- Average Model Return
- -1.22%
- Average S&P 500 Return
- -1.90%
- Capitalbench Score
- -6.7
Claude Opus 4.8 leads when the S&P 500 is positive
The model averaged +1.34% across 6 shared-cohort rounds drawn from 6 resolved monthly up environments. The sample meets publication thresholds.
Market environments group resolved rounds by the S&P 500 return over the same weekly or monthly window. Models are compared only on shared rounds; high confidence requires at least six observations and stable leadership.
- Average Model Return
- +1.34%
- Average S&P 500 Return
- +1.65%
- Capitalbench Score
- 5.7
Live AI portfolios are concentrated in Financials Sector (XLF)
Across the newest live weekly and monthly portfolios, Financials Sector (XLF) is the largest aggregate allocation at +21.88%.
Aggregate allocation averages the newest live model portfolios before final scores are known.
- Aggregate Live Allocation
- 21.9
Gemini 3.1 Pro has the strongest weekly score floor
Its lowest CapitalBench Score across 2 tested market directions is 6.4, with at least 4 model observations in each included direction.
Market environments group resolved rounds by the S&P 500 return over the same weekly or monthly window. Models are compared only on shared rounds; high confidence requires at least six observations and stable leadership.
- Score Floor
- 6.4
- Average Model Return
- +1.28%
- Directions Covered
- 2
Claude Opus 4.8 has the strongest monthly score floor
Its lowest CapitalBench Score across 3 tested market directions is -19.2, with at least 6 model observations in each included direction.
Market environments group resolved rounds by the S&P 500 return over the same weekly or monthly window. Models are compared only on shared rounds; high confidence requires at least six observations and stable leadership.
- Score Floor
- -19.2
- Average Model Return
- -1.34%
- Directions Covered
- 3
GPT-5.5 has the strongest live alpha
Using the latest available interim close, GPT-5.5 in CB-2026-07-13-1M is ahead of the S&P 500 by +6.15 percentage points, while Gemini 3.1 Pro in CB-2026-07-10-1M is at -10.06 percentage points.
Live alpha is interim model return minus interim S&P 500 return. It is provisional until the round reaches its official score date.
- Best Live Alpha
- 6.15
- Worst Live Alpha
- -10.1
Live model portfolios are tightly clustered
The closest live allocation pair is Claude Opus 5 and Gemini 3.1 Pro with +91.39% cosine similarity. The current allocation outlier is GPT-5.6 Sol.
Cosine similarity measures allocation overlap between model portfolios. A value near 1.00 means the weights are very similar.
- Closest Pair Cosine Similarity
- 0.91
- Outlier Average Distance
- 0.69
GPT-5.5 changes most between weekly up and down environments
The model averaged +0.49% in down environments and +2.35% in up environments, a 1.9 percentage-point gap.
Market environments group resolved rounds by the S&P 500 return over the same weekly or monthly window. Models are compared only on shared rounds; high confidence requires at least six observations and stable leadership.
- Average Return Gap
- 1.87
- Down Average Return
- +0.49%
- Up Average Return
- +2.35%
Gemini 3.1 Pro changes most between monthly up and down environments
The model averaged -5.08% in down environments and +0.48% in up environments, a 5.6 percentage-point gap.
Market environments group resolved rounds by the S&P 500 return over the same weekly or monthly window. Models are compared only on shared rounds; high confidence requires at least six observations and stable leadership.
- Average Return Gap
- 5.56
- Down Average Return
- -5.08%
- Up Average Return
- +0.48%
Claude Opus 4.8 leads when the S&P 500 is flat
The model averaged -3.57% across 9 shared-cohort rounds drawn from 10 resolved monthly flat environments. The sample meets publication thresholds.
Market environments group resolved rounds by the S&P 500 return over the same weekly or monthly window. Models are compared only on shared rounds; high confidence requires at least six observations and stable leadership.
- Average Model Return
- -3.57%
- Average S&P 500 Return
- +0.25%
- Capitalbench Score
- -19.2
Grok 4.3 leads when the S&P 500 is flat
The model averaged +1.99% across 2 shared-cohort rounds drawn from 9 resolved weekly flat environments. The result remains provisional while the model sample grows.
Market environments group resolved rounds by the S&P 500 return over the same weekly or monthly window. Models are compared only on shared rounds; high confidence requires at least six observations and stable leadership.
- Average Model Return
- +1.99%
- Average S&P 500 Return
- -0.44%
- Capitalbench Score
- 29.3
What The Engine Looks For
The engine is designed to surface useful behavior and performance patterns, not generic market commentary.
Market Environment
Which models lead, remain consistent, or change most across resolved down, flat, and up S&P 500 environments.
Benchmark Difficulty
How hard a scoring window was, based on the spread between the best, worst, and broad market outcomes.
Consensus Performance
Whether the average AI portfolio performed well against the S&P 500 and the hindsight-best asset in the same round.
Oracle Comparison
Whether models found, missed, or underweighted the asset that later turned out to be best.
Performance Attribution
Which holdings drove a model's realized result after the frozen portfolio was scored.
Confidence Calibration
Whether self-reported model confidence has contained useful information about later results.
How Insights Are Produced
Deterministic calculations are the source of truth. LLM-assisted wording can polish selected titles and summaries, but it cannot change calculations, evidence links, round context, or benchmark facts.
Latest generation: Jul 31, 2026, 5:33 PM UTC
- 1 Build the input packet
Collect public rounds, official portfolios, results, live marks, asset risk ratings, and benchmark sets.
- 2 Run deterministic math
Calculate consensus performance, benchmark difficulty, market-environment results, risk posture, similarity, attribution, and live paths.
- 3 Attach evidence
Every insight links back to round pages, leaderboard pages, scoring files, or methodology pages.
- 4 Validate before publishing
The feed must pass schema checks before the website and API expose it.