Choose two groups

Which results do you want to compare?

A round is one timed market test. A model group uses only rounds completed by every model being ranked together.

Bottom line

What changed?

Grok 4.3 ranks first in both groups. The groups have no completed rounds in common. Jul 24 Weekly includes 8 more rounds, while Aug 13 Weekly includes 2 more rounds. Grok 4.6 appears only in Aug 13 Weekly. GPT-5.5 appears only in Jul 24 Weekly.

Models in both groups71 only in Aug 13 Weekly; 1 only in Jul 24 Weekly
Rounds used by both0Jul 24 Weekly has 8 more; Aug 13 Weekly has 2 more
Did the order change?Changed a littleBased on models in both groups
Did the top model change?No2 of 3 stayed in the top three
Model by model

How did each model's result change?

Each group uses all of its completed rounds. Models that appear in only one group are marked clearly.

Jul 24 Weekly#1Grok 4.3xAIAug 13 Weekly#1No change
Jul 24 Weekly#3Claude Opus 5AnthropicAug 13 Weekly#2Up 1
Jul 24 Weekly#6Claude Fable 5AnthropicAug 13 Weekly#3Up 3
Jul 24 Weekly#4GPT-5.6 SolOpenAIAug 13 Weekly#4No change
Jul 24 Weekly#2Grok 4.5xAIAug 13 Weekly#5Down 3
Jul 24 Weekly#8Gemini 3.1 ProGoogleAug 13 Weekly#6Up 2
Jul 24 Weekly#5Claude Opus 4.8AnthropicAug 13 Weekly#7Down 2
Jul 24 Weekly#7GPT-5.5Only in Jul 24 WeeklyAug 13 WeeklyNot includedNot included
Jul 24 WeeklyNot includedGrok 4.6Only in Aug 13 WeeklyAug 13 Weekly#8Added
Fair comparison

Did both groups use the same rounds?

0 completed rounds are used by both groups. Jul 24 Weekly also uses 8 other rounds. Aug 13 Weekly also uses 2 other rounds. Different rounds can change scores and ranks.

0 used by both8 only in Jul 24 Weekly2 only in Aug 13 Weekly
Grok 4.3Overall score on 0 shared roundsn/aOverall score on 8 other Jul 24 Weekly rounds11.5Overall score on all 8 Jul 24 Weekly rounds11.5
Claude Opus 5Overall score on 0 shared roundsn/aOverall score on 8 other Jul 24 Weekly rounds9.8Overall score on all 8 Jul 24 Weekly rounds9.8
Claude Fable 5Overall score on 0 shared roundsn/aOverall score on 8 other Jul 24 Weekly rounds6.2Overall score on all 8 Jul 24 Weekly rounds6.2
GPT-5.6 SolOverall score on 0 shared roundsn/aOverall score on 8 other Jul 24 Weekly rounds8.8Overall score on all 8 Jul 24 Weekly rounds8.8
Grok 4.5Overall score on 0 shared roundsn/aOverall score on 8 other Jul 24 Weekly rounds10.8Overall score on all 8 Jul 24 Weekly rounds10.8
Gemini 3.1 ProOverall score on 0 shared roundsn/aOverall score on 8 other Jul 24 Weekly rounds2.3Overall score on all 8 Jul 24 Weekly rounds2.3
Claude Opus 4.8Overall score on 0 shared roundsn/aOverall score on 8 other Jul 24 Weekly rounds7.2Overall score on all 8 Jul 24 Weekly rounds7.2
See exactly which rounds were used

Used by both: None

Only in Jul 24 Weekly: CB-2026-07-24-1W, CB-2026-07-27-1W, CB-2026-07-28-1W, CB-2026-07-29-1W, CB-2026-07-30-1W, CB-2026-07-31-1W, CB-2026-08-04-1W, CB-2026-08-05-1W

Only in Aug 13 Weekly: CB-2026-08-15-1W, CB-2026-08-18-1W

Were any Aug 13 Weekly rounds left out?

No. Every possible Aug 13 Weekly round had a result from every model in the group.

What this means

Which results should you rely on?

Use Jul 24 Weekly as the more reliable ranking because it has 8 completed rounds. Aug 13 Weekly has 2 and needs 4 more before it has enough evidence to become the main ranking.

All results

What are all the numbers?

ModelIncluded inJul 24 Weekly rankAug 13 Weekly rankChangeJul 24 Weekly overall scoreAug 13 Weekly overall scoreOverall score change
Grok 4.3xAIBoth groups#1#1No change11.59.6-1.9
Claude Opus 5AnthropicBoth groups#3#2Up 19.86.2-3.6
Claude Fable 5AnthropicBoth groups#6#3Up 36.25.2-1.0
GPT-5.6 SolOpenAIBoth groups#4#4No change8.84.7-4.1
Grok 4.5xAIBoth groups#2#5Down 310.81.1-9.7
Gemini 3.1 ProGoogleBoth groups#8#6Up 22.30.3-2.1
Claude Opus 4.8AnthropicBoth groups#5#7Down 27.2-2.4-9.6
GPT-5.5OpenAIOnly Jul 24 Weekly#7Not included4.4n/an/a
Grok 4.6xAIOnly Aug 13 Weekly#8Addedn/a-4.5n/a