SportsHack.ai

Reference

Every term on this site, defined

This site publishes technical numbers, and technical numbers that nobody can check are just decoration. Everything below is defined precisely rather than simply — where a term has a real statistical meaning, that meaning is given, because a definition that is easy to read and wrong would undermine the point of publishing the numbers at all.

26 terms. Figures cited as ours come from the record and the backtest — measured, never targets.

Calibration

is the model honest about its own uncertainty

Brier score

For each yes/no event, take the difference between the stated probability and the outcome (1 or 0), square it, and average. Unlike ECE it penalises vagueness: always saying 50% scores 0.25 no matter what happens. That makes it the better single summary, because it can only be improved by being both well calibrated AND willing to move away from 50% when the evidence supports it. A Brier score means little on its own — it must be compared against a baseline for the same events.

lower is better. 0 is perfect; 0.25 is what you get by saying 50% to everything. What counts as good depends entirely on how predictable the events are.

Ours: 0.2375 on held-out 2025 strikeout starts, against 0.2806 for a naive baseline on the same starts.

Brier skill score

One minus the ratio of the forecast's Brier score to a baseline's. 0 means the forecast is no better than knowing how often the event happens in general; 1 would be perfection; negative means worse than the baseline. It is the honest way to report a Brier score, because it forces the comparison rather than leaving the reader to guess whether 0.24 is good.

higher is better.

Ours: +0.218 on Shawn's 2025 season backtest (0.1367 against a 0.1747 climatology baseline, 3,955 player-weeks).

Calibration

A forecast is calibrated when its stated probabilities match observed frequencies: of all the props given a 60% chance, close to 60% should hit. Calibration is about honesty, not sharpness — a model that says 50% about everything is perfectly calibrated and completely useless. It is also NOT the same as having an edge, because a forecast can match the market exactly and still be well calibrated.

closer to the stated value is better.

ECEExpected Calibration Error

Predictions are bucketed by stated probability, then each bucket's average prediction is compared with the share of outcomes that actually occurred. ECE is the average of those gaps, weighted by how many predictions fall in each bucket. It answers 'when this model says 70%, how far off is it really?' in a single number.

lower is better. Below ~0.02 is tight for a sample of a few thousand. Note that ECE shrinks toward zero on small samples for arithmetic reasons, so it is only meaningful next to its n.

Ours: 0.0049 on the strikeout model, measured on 4,570 starts from 2025 that the model never trained on.

Holdout

A model scored on the same data it learned from will flatter itself — it can memorise rather than generalise. A holdout is a slice of history withheld entirely from fitting, so scoring on it approximates how the model will do on games that have not happened yet. Every model on this site had to clear its gate on a holdout before it was allowed to serve.

Ours: The entire 2025 season is held out for the MLB models — 4,570 starts for strikeouts, 41,881 batter-games for hits, total bases and home runs.

Interval coverage

If a projection states a 50% interval, then about half of real outcomes should land inside it. Higher is NOT better here — 90% coverage on a stated 50% interval means the interval is far too wide to be useful. Closer to the stated number is better, in both directions.

closer to the stated value is better.

Ours: 48.6%–51.6% observed on stated 50% intervals across positions in the 2025 backtest.

PITProbability Integral Transform

For each prediction, PIT asks where the actual outcome landed inside the model's predicted distribution, as a value between 0 and 1. If the distribution is correct, those values are spread evenly across 0–1 — outcomes should land in the bottom 10% of the forecast about 10% of the time. Clustering near the edges means the model's intervals are too narrow (overconfident); clustering in the middle means too wide. It catches errors a single-number check cannot: a model can have the right average and still be badly wrong about its own uncertainty.

closer to the stated value is better. Mean PIT should sit near 0.50. Drift toward 0 or 1 is directional bias — the model is systematically projecting too high or too low.

Ours: 0.4977 mean PIT on the strikeout model's 2025 holdout.

Reliability band

A table version of calibration: all predictions between, say, 60% and 70% are pooled, and the band's average prediction is set against the share that actually happened. It shows WHERE a model is off, not just that it is — many models are well calibrated in the middle and overconfident at the extremes.

Walk-forward

The model is fitted using only information available before each game, then scored on that game, then the window rolls forward. This mirrors how the model is actually used and rules out lookahead — the subtle error where a test accidentally uses information that did not exist yet, which makes almost any model look brilliant.

Market math

what the prices mean

Closing line

Widely treated as the sharpest number available, because it reflects every pick, injury report and lineup card that arrived before the gate closed. It is the standard benchmark for judging whether an earlier price was good. On this site the closing snapshot is derived from real first-pitch times — the last capture strictly before the game began — and never from a label applied when the data was collected.

CLVClosing Line Value

The difference between the price available when a number was published and the closing price on that same side. Positive CLV means the market moved our way — others pushed the price in the direction we had already taken, which is evidence our number carried information the market had not yet absorbed. It is NOT a claim of profit: CLV can be positive on a prop that loses, and a single result says nothing either way. It is considered the earliest reliable signal of edge precisely because it does not depend on outcomes, which are noisy.

higher is better.

Ours: Not yet measurable. We publish at the same moment as our last line capture, so there is no later price to compare against — CLV is recorded as null rather than zero on all 629 settled props to date. A late capture now runs at 23:38 UTC specifically to create that second observation.

De-vigging

Both sides' implied probabilities are rescaled so they sum to 100%. We use the standard two-way method: P(over) divided by the sum of both sides' implied probabilities. This is an assumption, not a fact — it attributes the margin proportionally to each side, which is the conventional choice but not the only defensible one. Every market number on this site is de-vigged before it is compared to a model number, because comparing a model probability to a vigged price would manufacture disagreement that is really just the book's margin.

Disagreement

We deliberately say 'disagreement' rather than 'edge', 'value' or 'EV'. Those words assert that the difference is profitable — that the model is right and the market is wrong. Disagreement asserts only what we can actually demonstrate: two numbers differ, and here is by how much. Whether that difference is worth anything is exactly what the settlement ledger is accruing evidence about, and until it has, the stronger word would be a claim we cannot support.

Hold

Computed from how far the implied probabilities exceed 100%. Hold is where combos get expensive: it compounds with every leg, so a six-leg ticket built from legs each carrying a modest margin can hand the book a far larger share than any single leg suggests. The combo pages show hold precisely so that compounding is visible rather than buried.

lower is better.

Hypothetical ROI

Return on the amount staked, using one unit per prop with no scaling of any kind, priced at the best captured price at the moment of publication. It is labelled hypothetical because it is: it assumes those prices were available in the size you wanted, ignores limits and account restrictions, and is measured over a sample far too small to project forward. It is a track record, not a recommendation, and nothing on this site tells anyone what to pick.

Implied probability

The reciprocal of decimal odds: a price of 2.00 implies 50%, 1.50 implies 66.7%. Implied probabilities taken straight off the board always sum to more than 100% across the possible outcomes, because the excess is the book's margin. That excess must be removed before comparing a price to a model number — see de-vigging.

Push

Only possible on whole-number lines: a line of 5.5 strikeouts cannot push, but a line of 5 does if the pitcher records exactly five. In our settlement ledger a push scores zero profit and is excluded from both the hit rate and the ROI denominator, since no money was ever at risk on it.

Settlement

For every prop where the model disagreed with the price, settlement records which side the model took, how it resolved against the number, and what a flat stake would have returned. Calibration asks whether stated probabilities are honest; settlement asks whether the disagreements win. A model can be excellent at one and hopeless at the other, which is why the two records are kept separate and never combined into a single accuracy figure.

SGPSame-Game Combo

Legs within a game move together — a pitcher going deep with a big strikeout total makes his opponents' hitting props less likely. Multiplying such legs as if they were independent produces a joint probability that is simply wrong, usually flatteringly so. This site refuses same-game combinations outright rather than pricing them with a correlation model it has not validated.

Vigalso called juice or the overround

If a book offered a fair coin flip at true odds, it would make nothing. Instead both sides are priced slightly short, so the implied probabilities sum to more than 100%. That surplus is the vig. It is the reason a player who wins exactly half of their coin flips still loses money over time.

Model terms

how the numbers are built

Eligibility floor

Below the floor, a projection would be almost entirely prior with the player's name attached, which reads as knowledge the model does not have. Those players are shown as outside model competence with the reason stated, rather than being given a number that looks like the others.

Ours: Pitchers below the floor render as 'insufficient history — outside model competence' with their actual batters-faced count shown.

Park factor

Coors Field increases offence; pitcher-friendly parks suppress it. Our factors are estimated per venue from historical data and shrunk toward neutral in proportion to how little data a park has, so an unusual season at one stadium does not become a permanent belief about it.

Ours: Applied to the hits, total-bases and home-run models (v1.1-park). All three were re-validated through the full gate after the layer was added.

Posterior

The output that actually gets published — prior plus this game's inputs (opponent, park, batting order slot, starter faced). When a projection is described as a posterior mean, that is the average of the full predicted distribution, not a single guess with the uncertainty discarded.

Prior

Typically a blend of league baseline and the player's own history. Priors matter most where data is thinnest: a rookie with 40 plate appearances is pulled strongly toward the league baseline, while an established regular is trusted much closer to his own record.

Shadow variant

It issues into its own ledger alongside production, so the two can be compared on identical games over time. Promotion is decided by that accumulated record — at least 14 nights — rather than by whichever looked better on the first good evening.

The gate

Walk-forward fitting, an entire held-out season, randomised PIT within buckets, ECE within tolerance for the sample size, mean PIT near 0.50, and a Brier score beating stated baselines. The thresholds are set before results are seen, and a model that fails does not ship. Outs and earned runs failed twice and have never served.

Why the wording is careful

Two distinctions on this page do real work. The first is disagreementversus "edge" or "EV" — we can show that our number differs from the market's, but claiming that difference is profitable requires evidence we are still accruing. The second is calibration versus settlement — whether stated probabilities are honest, and whether the disagreements win, are different questions with different answers. They are reported separately and never averaged into one number.