Living equal-run rosters

All Benchmark Sets

Older sets keep accumulating shared rounds. New official rosters are opened automatically only when no existing set already contains the models in that run. Sets qualify at 6 weekly shared rounds or 3 monthly shared rounds.

Weekly Current Benchmark
Jul 10, 2026 roster

Weekly comparison set automatically opened when the Jul 10 official roster first required a new equal-run benchmark group across 8 models.

6 shared rounds 8 models Met current threshold
Open set
Monthly Current Benchmark
May 28, 2026 roster

Monthly comparison set that adds Claude Opus 4.8 to the established model roster.

21 shared rounds 5 models Met current threshold
Open set
Weekly Qualified
Jul 8, 2026 roster

Weekly comparison set that starts when Grok 4.5 joins the weekly benchmark roster.

8 shared rounds 7 models Met 6+ threshold
Open set
Weekly Qualified
Jun 9, 2026 roster

Weekly comparison set that starts when Claude Fable 5 joins the weekly benchmark roster.

15 shared rounds 6 models Met 6+ threshold
Open set
Weekly Qualified
May 28, 2026 roster

Weekly comparison set that adds Claude Opus 4.8 to the established model roster.

33 shared rounds 5 models Met 6+ threshold
Open set
Weekly Qualified
May 24, 2026 roster

Original weekly comparison set for the first four CapitalBench models.

35 shared rounds 4 models Met 6+ threshold
Open set
Monthly Qualified
May 10, 2026 roster

Original monthly comparison set for the first four CapitalBench models.

24 shared rounds 4 models Met 3+ threshold
Open set
Weekly Forming
Jul 21, 2026 roster

Weekly comparison set automatically opened when the Jul 21 official roster first required a new equal-run benchmark group across 7 models.

3 shared rounds 7 models 3/6 3 more to qualify
Open set
Monthly Forming
Jun 9, 2026 roster

Monthly comparison set that starts when Claude Fable 5 joins the monthly benchmark roster.

2 shared rounds 6 models 2/3 1 more to qualify
Open set
Weekly Waiting
Jul 24, 2026 roster

Weekly comparison set automatically opened when the Jul 24 official roster first required a new equal-run benchmark group across 8 models.

0 shared rounds 8 models 0/6 6 more to qualify
Open set
Monthly Waiting
Jul 24, 2026 roster

Monthly comparison set automatically opened when the Jul 24 official roster first required a new equal-run benchmark group across 8 models.

0 shared rounds 8 models 0/3 3 more to qualify
Open set
Monthly Waiting
Jul 21, 2026 roster

Monthly comparison set automatically opened when the Jul 21 official roster first required a new equal-run benchmark group across 7 models.

0 shared rounds 7 models 0/3 3 more to qualify
Open set
Monthly Waiting
Jul 10, 2026 roster

Monthly comparison set automatically opened when the Jul 10 official roster first required a new equal-run benchmark group across 8 models.

0 shared rounds 8 models 0/3 3 more to qualify
Open set
Monthly Waiting
Jul 8, 2026 roster

Monthly comparison set that starts when Grok 4.5 joins the monthly benchmark roster.

0 shared rounds 7 models 0/3 3 more to qualify
Open set
Selection rule

How Does a Set Become Current?

Fixed roster A set starts when the model roster changes. Missed rounds excluded If one set model is missing, that round is excluded for everyone. Newest qualified wins The newest qualified weekly or monthly set becomes current automatically.
Cross-set evidence

How do results change between model groups?

Main published results compared with the newest groups that have results.

Weekly Jul 10 Weekly vs Jul 21 Weekly

Grok 4.3 ranks first in Jul 10 Weekly. GPT-5.6 Sol ranks first in Jul 21 Weekly. The groups have no completed rounds in common. Jul 10 Weekly includes 6 more rounds, while Jul 21 Weekly includes 3 more rounds. Claude Opus 4.7 appears only in Jul 10 Weekly.

7models in both 0rounds used by both Changed a lotchange in order Yestop model changed

Use Jul 10 Weekly as the more reliable ranking because it has 6 completed rounds. Jul 21 Weekly has 3 and needs 3 more before it has enough evidence to become the main ranking.

Compare weekly groups
Monthly May 28 Monthly vs Jun 9 Monthly

Claude Opus 4.8 ranks first in May 28 Monthly. Grok 4.3 ranks first in Jun 9 Monthly. The groups share 2 completed rounds. May 28 Monthly includes 19 more rounds. Claude Fable 5 appears only in Jun 9 Monthly.

5models in both 2rounds used by both Changed a littlechange in order Yestop model changed

Use May 28 Monthly as the more reliable ranking because it has 21 completed rounds. Jun 9 Monthly has 2 and needs 1 more before it has enough evidence to become the main ranking.

Compare monthly groups