Working tool

Playoff odds and variance

Fantasy is a fourteen-game season decided head-to-head. Put your team's weekly average and spread into the model and see how often a better team still misses.

A fantasy season is short, and the scoring format throws away most of the information in it. You do not accumulate points against the field; you win or lose a single head-to-head each week. Score 140 against an opponent's 139 and it counts exactly as much as scoring 140 against 60.

That structure converts a modest scoring advantage into a coin flip with a slight lean, and then asks you to survive thirteen or fourteen of them.

Win chance per week
Expected wins
Chance of qualifying
If you add 5 pts/week

What the model does

Two steps, both standard.

Per-game win probability. Treat your weekly score and your opponent's as independent draws from normal distributions. The difference between two normals is itself normal, with the means subtracted and the variances added:

spread of the difference = square root of (your SD squared + their SD squared)
win chance = normal CDF of (your average − their average) divided by that spread

Season odds. Treat the games as independent and use the binomial distribution to get the chance of reaching your win target.

Both simplifications are visible and worth naming. Real leagues have unequal schedules, so your opponents are not all the league average. Games are not perfectly independent, since the same roster plays every week and a good or bad roster persists. Ties are ignored. None of that changes the conclusion, because the conclusion is driven by the size of the spread relative to the size of the edge, and that ratio is not delicate.

The uncomfortable arithmetic

Set both teams to identical averages and identical spreads and the model gives a 50% weekly win chance, which is correct and unsurprising. Now ask how often a pure coin-flip team wins eight of fourteen games. The answer is about 40% — and by symmetry, about 40% of coin-flip teams win six or fewer. Roughly a fifth land exactly on seven.

That is before you introduce any actual difference in team quality. Add a real edge and it improves, but not as fast as intuition suggests, because weekly scores swing by far more than the gap between a good roster and an average one. A weekly standard deviation in the mid-twenties against a scoring edge of four or five points a week means the noise is roughly five times the signal, every single week.

What this should change

Judge process, not outcome. A season is far too small a sample to grade a draft. If you want feedback that arrives faster than the noise, look at whether your starting lineups were the right ones given what you knew, and whether your projections were systematically off in a direction you could fix.

Variance is a lever, not a fixed cost. If you are the strongest team in the league, you want a lower spread: consistent producers, safe starting slots, no all-or-nothing lineup gambles. If you are behind and need to catch up, you want the opposite, because a wide distribution gives you more of the outcomes where you win a game you had no business winning. This is the single most actionable thing the model produces, and it points in opposite directions for different teams in the same league.

The regular season is a filter, not the prize. Most formats then run a single-elimination playoff of two or three rounds, which is a coin-flip tournament stapled to the end of a coin-flip season. Winning it is a genuine achievement in the sense that someone has to, and close to no evidence about roster quality.

Format changes the noise. Leagues that award playoff spots on total points as well as record are deliberately reducing variance. Two-week playoff matchups do the same. Both are defensible design choices, and both make the season more informative — worth knowing when you are the one voting on the rules.

A note on estimating your own inputs

The calculator is only as good as the four numbers you feed it, and the spreads are the ones people guess worst. Your weekly standard deviation is not the range between your best and worst week; it is roughly the typical distance from your own average, which in most formats sits somewhere in the twenties for a standard starting lineup. Scoring settings move it: more scoring events at smaller values — receptions, yardage — reduce it, while touchdown-heavy systems inflate it substantially, because touchdowns are lumpy and rare.

If you want a rough figure from your own history, take a season of your weekly scores, find the average, and eyeball how far a typical week sits from it. That is close enough for this purpose, and closer than the number most people would have guessed.