METHOD / VERSION nerf-watch-analysis-v1
What counts as
a real decline?
A worse score needs comparable tests, a meaningful effect, and an independent repeat.
What we measure
The first release imports Kingy.ai’s fixed bank of 12 Python function tasks and 12 quantitative JSON-answer tasks. Each model attempts each task three times. A pass requires all frozen deterministic checks; individual assertions are not separate samples. We neither execute imported outputs nor use an LLM judge.
Comparable contracts
Task-bank, protocol, model-settings, grader, and category hashes define a series. Changed prompts, fixtures, scoring, provider routes, settings, or identities separate the series. The disclosed v1 offline correction preserves original grades. v2 starts a new baseline.
Uncertainty bands
We use a reproducible task-cluster bootstrap with 12,000 resamples. Tasks are sampled as clusters, then retained repetitions are resampled within each task. Tasks receive equal weight. Plot bars are descriptive 95% intervals of task/repeat variation. With only three repeats, these bands have limited resolution and cannot establish temporal stability or predict an unsampled endpoint.
Weekly comparisons
Weeks run Monday through Sunday in America/Vancouver. Reports target the last completed week and its preceding week. We select the earliest complete cohort in the target week and the last complete comparable cohort in the prior week; later cohorts may serve as independent confirmation, including in the current reporting week. This predeclared rule avoids selecting the worst run. Sparse weekly sampling does not cover every day; missing weeks remain gaps.
Matched comparisons resample task pairs, preserving task identity across cohorts. We use a 99.167% interval for each of six primary model/category comparisons (Bonferroni family alpha 5%). Five percentage points is the minimum meaningful change. The rules are fixed before new Nerf Watch comparisons are collected.
Status rules
“Possible decline” means the point estimate dropped at least five percentage points; the interval and confirmation status remain visible. “Decline reproduced” requires the entire adjusted interval below −5 points in the primary cohort and a separate matching cohort at least 24 hours later. “Improvement” requires the interval above +5 points. “No detectable change” means these criteria were not met; it does not prove equivalence. Incomplete coverage or incompatible/missing weeks yields “Insufficient evidence.”
Service failures and cost
Responses, failed attempts, refusals, output-cap stops, and latency are displayed separately. Settled failed attempts remain in the attempted-task denominator. Unresolved accounting or incomplete source cohorts cannot establish a trend. Refusal counts use explicit source grading flags; an unflagged refusal may be a failed task. Costs are estimated until provider billing is reconciled. We do not claim a cost winner.
Refreshes and collection
The report targets a Monday 09:00 Pacific refresh through the existing hosted monitoring lane. A refresh imports available evidence and creates a bundle; it never makes a model request or changes its test date. Configuration and ordinary scheduled-run proof are separate. The original feature’s funding is never used for Nerf Watch collection.
Evidence and archives
Every imported attempt retains its exact prompt, output, original response hash, request identifier where exposed, settings, grading outcome, timestamps, usage, and cost status. Weekly bundles include checksums and retained grader sources. Private hidden fixtures and transport envelopes remain with the source owner and are explicitly withheld. A pending bundle has no DOI; only a verified production archive can supply one.