SportsHack.ai
Menu

Retrospective backtest

Scored after the fact against completed 2025 games. This is NOT the live forward ledger and must never be presented as one. For projections graded as they happen, with nothing known in advance, see the live forward record.

2025 season, scored after the fact

Shawn Childs' weekly projections for the complete 2025 season, graded against what actually happened. Because the games were already played when this was computed, it answers "how good were these numbers?" — not "how good are we live."

18 weekly TeamsSplits workbooks (weeks 1-18; week 17 was initially absent and supplied 2026-08-04) x the validated 2025 player-game ingest

3,955
player-weeks scored
0.1367
Brier score
0.218
skill vs base rate
5.2 pts
mean calibration error

Brier score measures probabilistic accuracy — lower is better, and 0 is perfect. Against a baseline that simply predicts each event's base rate (0.1747), these projections score 0.1367, a skill score of 0.218. Positive skill means the projections carried real information beyond knowing how often the event happens in general.

Expected vs observed

Every predicted probability, bucketed, against how often the event actually occurred. A perfectly calibrated forecast puts observed equal to predicted in every row.

Predicted bandAvg predictedObservedGapn
0–10%4.2%4.2%+0.04,867
10–20%14.4%14.2%-0.21,962
20–30%24.3%22.2%-2.11,016
30–40%35.0%31.7%-3.3694
40–50%44.5%38.0%-6.5458
50–60%55.5%36.8%-18.7628
60–70%64.8%47.3%-17.5640
70–80%74.6%58.7%-15.9618
80–90%85.1%61.8%-23.3476
90–100%94.8%75.9%-18.9506

By position, with interval coverage

Coverage asks whether the uncertainty is honest: of the outcomes that should land inside a 50% interval, how many did? A well-calibrated interval covers close to its stated rate — higher is not better here, closer is.

PositionPlayer-weeksMAE (FP)Brier50% interval80% interval
QB5255.780.170949.7%81.1%
RB1,0815.010.130849.0%78.2%
WR1,6295.090.127051.6%81.1%
TE7204.340.142648.6%78.9%

By sample size

Absolute error rises with sample size because players with more prior weeks are higher-usage and carry larger projections. Relative error (MAE / mean projection) is the like-for-like comparison and moves the other way.

Prior weeks on that playernMean projectionMAERelative MAE
0–2 prior weeks5303.993.190.80
3–5 prior weeks1,0979.524.790.50
6–9 prior weeks1,15610.655.370.50
10–17 prior weeks1,17211.725.730.49

Week by week

WeekPlayer-weeksMAEBrier
Week 42884.760.1315
Week 52514.960.1425
Week 62554.880.1326
Week 72695.190.1358
Week 82334.900.1433
Week 92445.180.1440
Week 102474.830.1309
Week 112665.400.1463
Week 122424.790.1305
Week 132764.920.1392
Week 142544.950.1406
Week 152755.330.1375
Week 162925.240.1432
Week 172854.920.1346
Week 182785.030.1191

Method

  • · Probabilities are derived empirically, walk-forward: for week N and position P, the residual pool is (actual - projected) from weeks before N only.
  • · P(actual >= threshold) = the share of that pool for which projected + residual clears the threshold. No distribution is assumed.
  • · Weeks 1-3 are burn-in that builds the first pool and are not scored.
  • · Predictive intervals use the same pool's quantiles; coverage is the share of outcomes falling inside.
  • · Scoring: PPR for RB/WR/TE, Shawn's 4-point passing-TD scheme for QB.
  • · Projected-but-did-not-play rows are excluded here and counted separately in the season report.

Weeks covered: 1–18 complete (weeks 4–18 scored after burn-in). Thresholds scored: QB 15/20/25 · RB 10/15/20 · WR 10/15/20 · TE 8/12/16 fantasy points. 11,865 threshold events in total.

This page is retrospective by construction and is kept separate from the live forward record on purpose. The two are never combined into a single claim.