---
title: "How BRADY Works — Held Out"
description: "How BRADY rates NFL quarterbacks, and the held-out numbers behind every part of it."
url: "https://heldout.org/methodology/brady"
---

Methodology 

# How BRADY Works

## What BRADY Measures

**BRADY** is _Bayesian Regressed Adjusted Down Yield_. Bayesian and regressed, because every part of it is pulled toward the league by however little is known about it. Adjusted, for the defences a quarterback faced. A yield per down: the expected points he produces for each play he runs.

BRADY estimates how much offensive value a quarterback is worth per play he runs, using only his earlier games. It is shown on the same scale as RICE, where **75 is the median starter** and the best estimate of a quarterback's ability on record — Peyton Manning's at the end of 2009 — reads 99, with Drew Brees's 2011 just behind. A single-season rating is shrunk less, so the best season on that rating reads higher: Manning's 2004 is 104.0.

A play is every dropback, every designed run and every sneak. Designed runs count as plays because their value is counted: a rate whose top line includes what a quarterback does with his legs has to have those carries in its bottom line too. Penalties that wipe out a snap are left out, because most of them — false starts, holding — belong to the line.

The centre is the median _starter_, not the average of everyone who has thrown a pass. Centring on the latter put the median starter above 80 and squeezed thirty real starting quarterbacks into the top fifth of the scale, which makes the number useless for the comparison people actually want.

The question it is built to answer is deliberately narrow: **given everything this quarterback has done so far, how good is he, in a way that should still be true two months from now?** Not how good his last few games looked. Not whether he deserves an award.

## Two BRADYs

BRADY is one estimator tuned for two jobs, and each version is graded only on its own job.

**Descriptive BRADY** is on every page of this site: the grades, the season and career boards, WAR, and the grade on each matchup card. Its job is to say how good a quarterback is, and it is graded the way that claim should be: on a team's next eight games, by the offence it produces and by how well the RICE rating this site shows predicts them with it; and on January, every playoff game and the champion, from ratings taken as the regular season ends. It is checked against the MVP and All-Pro vote.

**Predictive BRADY** is the one inside RICE. Its job is the next few Sundays, and it is graded on exactly that: how well RICE predicts each team's next three games with it, the offence those games produce, and whether it gets the size of the fall right when a starter is hurt and a backup comes in.

They differ in three settings, each chosen on the seasons before 2020 and checked once on the seasons since. The descriptive rating's shortest memory is four games where the predictive one's is two, because a claim about a career should not swing on one bad afternoon — and every shorter memory predicts the next eight games worse. It weighs each defence's most recent games a little more heavily. And it credits the quarterback sneak, which the predictive rating leaves out because the backup test below found the offensive line responsible for most of it.

On ninety-eight ranked seasons in a hundred the two are within two points of each other, and they are never more than three apart. Where they differ, the one on this page is the one to trust about a career, and the one inside RICE is the one to trust about Sunday.

## How It Works

**Every play is broken into parts.** A completion, an incompletion, a sack, a scramble, an interception, a pass-interference flag drawn, a designed run, a sneak and a lost fumble are each valued separately, in expected points.

The reason for taking it apart is that the parts differ enormously in how much they tell you about the quarterback. Interceptions are real but arrive so rarely that a single season of them is close to noise. Who recovers a fumble is close to a coin flip, so a lost fumble is its own part rather than muddying the sack or scramble it happened on. Splitting a completion into the throw and the run after the catch was tested and told the rating nothing new, so a completion is carried whole.

**Each part is adjusted for the defences he faced.** Separately, because a defence that is good at pressuring the quarterback is not necessarily good at covering receivers.

**Each part is then shrunk toward average by how reliably it is measured.** A component that is mostly noise gets pulled hard toward the league mean; one that is mostly signal is left nearly alone. This is the part that makes the rating stable without making it stubborn.

**The whole thing is averaged over four different memories**, from four games to thirty-two. How quickly a quarterback changes is not one number — a 23-year-old in his second season and a 38-year-old are different problems — so rather than pick a rate of forgetting, the rating averages over a spread of them.

## Career and Season-Only

**Season-only** starts each September from scratch and learns from that year alone. It is the default for a finished season, and it agrees with the AP ballot considerably better: the average miss against MVP and All-Pro voting is 1.8 places, against 2.2 for the career rating.

**Career** carries everything a quarterback has ever done, decayed toward the present. It is the better guide to what comes next, and it is the default while a season is still being played, because three games is not enough to rate anybody on their own.

## Why It Disagrees With Award Voters

It usually has a good reason to and sometimes does not.

BRADY is a career rating. A quarterback's first outstanding season is pulled toward a mediocre past, because from the model's point of view one good year is weak evidence about a player with three bad ones. Awards are for the season alone. These are different questions and the answers should differ.

Shrinkage also pulls extremes toward the middle by design, and awards are given to extremes. A rating built to still be true next season will always give some of this away. That is a statement about what the rating is for, not a defect waiting to be tuned out.

Where a quarterback won an award, it is marked in the tables. Where the rating puts him fourth, that disagreement is shown rather than smoothed over.

## What It Deliberately Gives Up

There are two ways to grade a quarterback rating, and they pull against each other.

One asks whether it predicts how much the offence will score. The other asks whether it is describing the _quarterback_ rather than the team around him.

A rating can do very well on the first by quietly measuring the supporting cast — a good offensive line and good receivers really do predict scoring — while getting worse at the thing it claims to measure. So BRADY is checked against a second test built on games where a starter got hurt and a backup took over: the roster is the same, only the passer changed.

Three signals that raised the first score were dropped because they failed the second. The clearest was a measure of how pass-heavy a team's play-calling was. It looked like one of the strongest inputs in the model. It was forecasting the scoreboard: a worse quarterback means a team behind, and a team behind throws more.

Removing it made the rating measurably worse at predicting offensive production, and measurably better at describing a quarterback. That trade is the reason for the whole exercise.

## Wins Above Replacement

BRADY says how good a quarterback is per play. It cannot say what he was _worth_ — being excellent for four games is not the same as being excellent for seventeen.

**WAR** counts the games. For each one he started, it asks how many more wins an otherwise average team would expect against an average opponent with him instead of a replacement-level quarterback. A game he left at half-time counts as half a game.

The conversion from rating to wins is the quarterback's whole contribution to his team. Across every team-game since 2003, a point of BRADY goes with 31 points a game of team strength, and a starter's BRADY accounts for about 40% of how good his team is — close to the third the research on the position arrives at. That relationship has not changed from one era to the next.

Each season is measured against itself: a quarterback is compared with that season's median starter, and a replacement sits a fixed distance below it. So a 2007 season and a 2023 season are judged by how far each stood above the quarterbacks of its own year, and a legendary season from the 2000s reads as one.

Seasons are compared on **WAR per game started**, so a sixteen-game season stands beside a seventeen-game one, and a season needs eight starts to be ranked. The career total — every game, added up — appears on one page only: the all-time career board, where longevity is the point.

It deliberately does **not** multiply by how often he threw. A quarterback's busiest games are his losses — teams throw when they are behind — so a WAR that counted dropbacks paid passers for trailing. And a passer who throws fifty times a game moves the result no more than one of the same rating who throws thirty: scaling the rating by volume makes RICE's predictions worse, not better.

Replacement is what quarterbacks who were second or lower on their team's depth chart actually produced in the games they started: **0.127 expected points a play** below the season's median starter. The obvious alternative — everyone who threw fewer than 250 times in a season — is contaminated, because it sweeps in every starter who got hurt in September, and those men are good.

The median ranked season is worth about 0.12 wins a start over a replacement, in every era. A very good one is 0.24, and the best on record — Drew Brees's 2011, Tom Brady's 2007 and Peyton Manning's 2004 — are about 0.36: some six wins over seventeen starts. Across a career, Tom Brady is 86 wins above replacement.

Every figure here is held out — fitted on earlier seasons, scored on games the model had not seen. The full record, failures included, is in `docs/`.

Last Updated: Sep 26, 2026, 3:56 PM PDT © 2026, Aditya Kishore. All Rights Reserved. Club marks are the clubs’ own trademarks. Past marks courtesy of [SportsLogoHistory.com](https://sportslogohistory.com/nfl-primary-logo/) and data courtesy of [nflverse](https://nflverse.nflverse.com/).
