CapitalBench Score
A score of 30 means the model earned 30% of the best possible return across these rounds. Calculation
A score of 30 means the model earned 30% of the best possible return across these rounds. Calculation
Weekly comparison set
Original weekly comparison set for the first four CapitalBench models.
Every ranked model in this set is scored only on rounds that all 4 listed models completed. If one model misses a resolved round, that round is excluded from this set for everyone.
Every ranked model in this set completed the same 35 weekly rounds.
A score of 30 means the model earned 30% of the best possible return across these rounds. Calculation
A score of 30 means the model earned 30% of the best possible return across these rounds. Calculation
Average portfolio return across the same finished rounds.
Grok 4.3 ranks first in both groups. The groups share 6 completed rounds. May 24 Weekly includes 29 more rounds. Grok 4.5, Claude Fable 5, GPT-5.6 Sol, Claude Opus 4.8 appear only in Jul 10 Weekly.
Jul 10 Weekly is the main published ranking. May 24 Weekly also has enough rounds, so compare them to see whether the results hold across different model groups.
Compare these groupsThis roster stays fixed so the set can keep growing as a clean equal-run comparison.
Included rounds count toward the score. Excluded rounds are resolved rounds after the set started where at least one set model was missing.
CapitalBench Score equals total model return across included shared rounds divided by total max-possible return across those same rounds, multiplied by 100. Max possible is the best eligible asset in each included round in hindsight.