Choose two groups

Which results do you want to compare?

A round is one timed market test. A model group uses only rounds completed by every model being ranked together.

Bottom line

What changed?

GPT-5.6 Sol ranks first in Jul 24 Weekly. Claude Fable 5 ranks first in Aug 19 Weekly. The groups have no completed rounds in common. Jul 24 Weekly includes 11 more rounds, while Aug 19 Weekly includes 3 more rounds. Grok 4.6 appears only in Aug 19 Weekly. GPT-5.5, Claude Opus 4.8 appear only in Jul 24 Weekly.

Models in both groups61 only in Aug 19 Weekly; 2 only in Jul 24 Weekly
Rounds used by both0Jul 24 Weekly has 11 more; Aug 19 Weekly has 3 more
Did the order change?Changed a lotBased on models in both groups
Did the top model change?Yes1 of 3 stayed in the top three
Model by model

How did each model's result change?

Each group uses all of its completed rounds. Models that appear in only one group are marked clearly.

Jul 24 Weekly#6Claude Fable 5AnthropicAug 19 Weekly#1Up 5
Jul 24 Weekly#4Claude Opus 5AnthropicAug 19 Weekly#2Up 2
Jul 24 Weekly#1GPT-5.6 SolOpenAIAug 19 Weekly#3Down 2
Jul 24 WeeklyNot includedGrok 4.6Only in Aug 19 WeeklyAug 19 Weekly#4Added
Jul 24 Weekly#8Gemini 3.1 ProGoogleAug 19 Weekly#5Up 3
Jul 24 Weekly#5GPT-5.5Only in Jul 24 WeeklyAug 19 WeeklyNot includedNot included
Jul 24 Weekly#2Grok 4.5xAIAug 19 Weekly#6Down 4
Jul 24 Weekly#7Claude Opus 4.8Only in Jul 24 WeeklyAug 19 WeeklyNot includedNot included
Jul 24 Weekly#3Grok 4.3xAIAug 19 Weekly#7Down 4
Fair comparison

Did both groups use the same rounds?

0 completed rounds are used by both groups. Jul 24 Weekly also uses 11 other rounds. Aug 19 Weekly also uses 3 other rounds. Different rounds can change scores and ranks.

0 used by both11 only in Jul 24 Weekly3 only in Aug 19 Weekly
Claude Fable 5Overall score on 0 shared roundsn/aOverall score on 11 other Jul 24 Weekly rounds8.4Overall score on all 11 Jul 24 Weekly rounds8.4
Claude Opus 5Overall score on 0 shared roundsn/aOverall score on 11 other Jul 24 Weekly rounds9.6Overall score on all 11 Jul 24 Weekly rounds9.6
GPT-5.6 SolOverall score on 0 shared roundsn/aOverall score on 11 other Jul 24 Weekly rounds13.8Overall score on all 11 Jul 24 Weekly rounds13.8
Gemini 3.1 ProOverall score on 0 shared roundsn/aOverall score on 11 other Jul 24 Weekly rounds2.9Overall score on all 11 Jul 24 Weekly rounds2.9
Grok 4.5Overall score on 0 shared roundsn/aOverall score on 11 other Jul 24 Weekly rounds12.8Overall score on all 11 Jul 24 Weekly rounds12.8
Grok 4.3Overall score on 0 shared roundsn/aOverall score on 11 other Jul 24 Weekly rounds11.4Overall score on all 11 Jul 24 Weekly rounds11.4
See exactly which rounds were used

Used by both: None

Only in Jul 24 Weekly: CB-2026-07-24-1W, CB-2026-07-27-1W, CB-2026-07-28-1W, CB-2026-07-29-1W, CB-2026-07-30-1W, CB-2026-07-31-1W, CB-2026-08-04-1W, CB-2026-08-05-1W, CB-2026-08-07-1W, CB-2026-08-09-1W, CB-2026-08-11-1W

Only in Aug 19 Weekly: CB-2026-08-19-1W, CB-2026-08-20-1W, CB-2026-08-21-1W

Were any Aug 19 Weekly rounds left out?

No. Every possible Aug 19 Weekly round had a result from every model in the group.

What this means

Which results should you rely on?

Use Jul 24 Weekly as the more reliable ranking because it has 11 completed rounds. Aug 19 Weekly has 3 and needs 3 more before it has enough evidence to become the main ranking.

All results

What are all the numbers?

ModelIncluded inJul 24 Weekly rankAug 19 Weekly rankChangeJul 24 Weekly overall scoreAug 19 Weekly overall scoreOverall score change
Claude Fable 5AnthropicBoth groups#6#1Up 58.428.2+19.8
Claude Opus 5AnthropicBoth groups#4#2Up 29.622.8+13.2
GPT-5.6 SolOpenAIBoth groups#1#3Down 213.818.9+5.1
Grok 4.6xAIOnly Aug 19 Weekly#4Addedn/a16.1n/a
Gemini 3.1 ProGoogleBoth groups#8#5Up 32.914.1+11.2
GPT-5.5OpenAIOnly Jul 24 Weekly#5Not included8.8n/an/a
Grok 4.5xAIBoth groups#2#6Down 412.811.5-1.3
Claude Opus 4.8AnthropicOnly Jul 24 Weekly#7Not included6.9n/an/a
Grok 4.3xAIBoth groups#3#7Down 411.49.3-2.1