← All articles

Red-zone efficiency is a small-sample stat

Few numbers get treated as a more durable team trait than red-zone efficiency — the share of trips inside the twenty that an offense turns into touchdowns rather than field goals. A team that ranks near the top is praised for finishing, for having the personnel and the play-caller to punch it in when the field shrinks; a team near the bottom is diagnosed with a red-zone problem that demands a new package, a goal-line back, or a different coordinator. The diagnosis assumes the number is measuring a stable skill. It mostly is not. Red-zone touchdown percentage is computed over a startlingly small number of snaps, it swings hard from season to season, and a large part of any year’s ranking is the kind of noise that tells you nothing about next week, let alone next year.

How few trips we are talking about

An NFL offense reaches the red zone something like three or four times a game, which adds up to roughly fifty to sixty trips across a full season. Touchdown percentage is then computed over that handful of opportunities. Fifty-odd binary outcomes is a tiny sample by any statistical standard, and tiny samples are dominated by variance: the difference between a top-five red-zone offense and a bottom-ten one can be a small number of trips that broke one way rather than the other. A dropped fade, a goal-line spot that goes against you, a defensive penalty that gifts a fresh set of downs — a few of those across a season move a team several places in a ranking that fans and analysts then treat as a settled fact about the offense.

The tell is in the year-to-year wobble

The clean way to test whether a stat measures a real, repeatable skill is to ask how well this year predicts next year. Stable skills persist; noise does not. Red-zone touchdown rate is one of the less stable team numbers in football: teams that finish at the top of the league one season routinely drift toward the middle the next, and bottom-dwellers bounce back up without changing much about how they play. That regression is the signature of a number driven heavily by variance rather than by a fixed underlying trait. If a metric won’t hold still from one year to the next, building a narrative on a single season’s value is building on sand.

The compressed field changes the math

There is a structural reason red-zone results are so jumpy, and it is geometry. Inside the twenty the field is short and gets shorter, which compresses the defense, removes the deep space an offense uses to create separation, and turns drives into a sequence of high-variance, contested plays in tight windows. Outcomes in that environment hinge on a single broken tackle, a coverage bust, a coin-flip spot, or a fingertip catch far more than outcomes in the open field do. The compression doesn’t just make scoring harder; it makes the results noisier, because small events swing whole possessions. So the part of the field that fans scrutinize most closely is precisely the part where the sample is thinnest and the luck is loudest.

Why the narrative outruns the evidence

Red-zone failures are vivid. A stalled drive that settles for three after first-and-goal is the kind of sequence that sticks in memory and dominates the broadcast, so a few of them across a season can cement a reputation that the full body of evidence doesn’t support. The eye sees the empty-handed trips and concludes the offense can’t finish, when the underlying rate may be ordinary and the visible failures may be the normal variance of a small sample landing in a clump. Coaches feel the same pull: a bad red-zone month invites a scheme overhaul aimed at a problem that may not exist, treating a run of variance as a structural flaw and sometimes breaking something that was working in the process.

What actually carries from week to week

The irony is that the thing that best predicts how many touchdowns a team will score is not its red-zone rate but how often it reaches the red zone in the first place. Getting there is a function of moving the ball — early-down efficiency, explosive plays, staying on schedule — and those skills are measured over hundreds of snaps a season, so they stabilize and carry forward. A team that reaches the twenty constantly will score plenty of touchdowns even at a league-average finishing rate, while a team that rarely gets there can post a gaudy red-zone percentage and still struggle to put up points. Volume of trips is the signal; conversion rate is mostly the noise riding on top of it.

A quick thought experiment makes the point. Imagine two offenses with identical, league-average finishing ability, one that reaches the red zone four times a game and one that reaches it twice. Over a season the first will score far more touchdowns than the second while posting the same conversion rate, and if the second offense happens to convert a few of its rare trips, it can even top the red-zone leaderboard while scoring fewer points. The ranking that fans read as a measure of finishing would have the worse offense ahead. That inversion is the tell: a stat that can rank the lower-scoring team higher is not capturing the thing — points — that everyone is actually trying to measure.

What to read instead

Read red-zone results the way you would read any small sample: with a heavy thumb on the regression scale. A single season’s touchdown rate should be assumed to be closer to league average than it looks, and a multi-year rate is far more trustworthy than one year. Better still, lead with the inputs that hold still — how often the offense reaches the red zone, its overall efficiency between the twenties, its explosive play rate — because those predict future scoring more reliably than the finishing percentage everyone quotes. If you want to evaluate finishing specifically, look at it over several seasons and against the difficulty of the looks, not as a single year’s headline.

The honest read

Red-zone efficiency is a real outcome — those touchdowns and field goals genuinely happened, and over a long enough horizon a persistently great or poor finishing team will reveal itself. But as it is normally used, a single season’s red-zone rate read as a fixed identity, it is mostly a small-sample stat dressed up as a trait. It is computed over a few dozen high-variance trips, it regresses hard year to year, and it gets outpredicted by the simple question of how often a team reaches the red zone at all. Treat a hot or cold red-zone season as provisional, weight it toward the mean, and put your confidence in the volume of trips — the part of the story that actually repeats.