kingy.wikiIndependent measurements

NERF WATCH / WEEKLY MODEL CHECK

Did your model
get worse this week?

Read the measurements, the uncertainty, and the receipts. We track coding and math performance without smoothing away failures or missing weeks.

Models under observation

Week beginning 2026-09-28

openai

GPT-6.1 Sol

Coding: Insufficient evidenceMath: Insufficient evidence
Latest coding cohort
100.0%
Latest math cohort
100.0%

Each category uses 12 tasks × 3 attempts. These are cohort scores; they are not a weekly decline verdict.

Measurements and receipts →

anthropic

Claude Sonnet 5.5

Coding: Insufficient evidenceMath: Insufficient evidence
Latest coding cohort
86.1%
Latest math cohort
5.6%

Each category uses 12 tasks × 3 attempts. These are cohort scores; they are not a weekly decline verdict.

Measurements and receipts →

google

Gemini 3.8 Flash

Coding: Insufficient evidenceMath: Insufficient evidence
Latest coding cohort
97.2%
Latest math cohort
88.9%

Each category uses 12 tasks × 3 attempts. These are cohort scores; they are not a weekly decline verdict.

Measurements and receipts →

Two baselines. Two separate series.

v1 was measured October 4, 2026 UTC. v2 was measured October 5, 2026 UTC after prompts, fixtures, and scoring changed. Comparing their scores cannot show that a model improved or declined. Both remain available with their original evidence and disclosed corrections.

Weekly reports compare adjacent completed weeks with the same measurement contract. The report compares 2026-09-21 and 2026-09-28. Model pages show whether each category has a comparable pair; a missing pair yields “Insufficient evidence.”