What BRADY Measures
BRADY is Bayesian Regressed Adjusted Down Yield. Bayesian and regressed, because every part of it is pulled toward the league by however little is known about it. Adjusted, for the defences a quarterback faced. A yield per down: the expected points he produces for each play he runs.
BRADY estimates how much offensive value a quarterback is worth per play he runs, using only his earlier games. It is shown on the same scale as RICE, where 75 is the median starter and the best estimate of a quarterback's ability on record — Peyton Manning's at the end of 2009 — reads 99, with Drew Brees's 2011 just behind. A single-season rating is shrunk less, so the best season on that rating reads higher: Manning's 2004 is 104.0.
A play is every dropback, every designed run and every sneak. Designed runs count as plays because their value is counted: a rate whose top line includes what a quarterback does with his legs has to have those carries in its bottom line too. Penalties that wipe out a snap are left out, because most of them — false starts, holding — belong to the line.
The centre is the median starter, not the average of everyone who has thrown a pass. Centring on the latter put the median starter above 80 and squeezed thirty real starting quarterbacks into the top fifth of the scale, which makes the number useless for the comparison people actually want.
The question it is built to answer is deliberately narrow: given everything this quarterback has done so far, how good is he, in a way that should still be true two months from now? Not how good his last few games looked. Not whether he deserves an award.
Two BRADYs
BRADY is one estimator tuned for two jobs, and each version is graded only on its own job.
Descriptive BRADY is on every page of this site: the grades, the season and career boards, WAR, and the grade on each matchup card. Its job is to say how good a quarterback is, and it is graded the way that claim should be: on a team's next eight games, by the offence it produces and by how well the RICE rating this site shows predicts them with it; and on January, every playoff game and the champion, from ratings taken as the regular season ends. It is checked against the MVP and All-Pro vote.
Predictive BRADY is the one inside RICE. Its job is the next few Sundays, and it is graded on exactly that: how well RICE predicts each team's next three games with it, the offence those games produce, and whether it gets the size of the fall right when a starter is hurt and a backup comes in.
They differ in three settings, each chosen on the seasons before 2020 and checked once on the seasons since. The descriptive rating's shortest memory is four games where the predictive one's is two, because a claim about a career should not swing on one bad afternoon — and every shorter memory predicts the next eight games worse. It weighs each defence's most recent games a little more heavily. And it credits the quarterback sneak, which the predictive rating leaves out because the backup test below found the offensive line responsible for most of it.
On ninety-eight ranked seasons in a hundred the two are within two points of each other, and they are never more than three apart. Where they differ, the one on this page is the one to trust about a career, and the one inside RICE is the one to trust about Sunday.
How It Works
Every play is broken into parts. A completion, an incompletion, a sack, a scramble, an interception, a pass-interference flag drawn, a designed run, a sneak and a lost fumble are each valued separately, in expected points.
The reason for taking it apart is that the parts differ enormously in how much they tell you about the quarterback. Interceptions are real but arrive so rarely that a single season of them is close to noise. Who recovers a fumble is close to a coin flip, so a lost fumble is its own part rather than muddying the sack or scramble it happened on. Splitting a completion into the throw and the run after the catch was tested and told the rating nothing new, so a completion is carried whole.
Each part is adjusted for the defences he faced. Separately, because a defence that is good at pressuring the quarterback is not necessarily good at covering receivers.
Each part is then shrunk toward average by how reliably it is measured. A component that is mostly noise gets pulled hard toward the league mean; one that is mostly signal is left nearly alone. This is the part that makes the rating stable without making it stubborn.
The whole thing is averaged over four different memories, from four games to thirty-two. How quickly a quarterback changes is not one number — a 23-year-old in his second season and a 38-year-old are different problems — so rather than pick a rate of forgetting, the rating averages over a spread of them.
Career and Season-Only
Season-only starts each September from scratch and learns from that year alone. It is the default for a finished season, and it agrees with the AP ballot considerably better: the average miss against MVP and All-Pro voting is 1.8 places, against 2.2 for the career rating.
Career carries everything a quarterback has ever done, decayed toward the present. It is the better guide to what comes next, and it is the default while a season is still being played, because three games is not enough to rate anybody on their own.
Why It Disagrees With Award Voters
It usually has a good reason to and sometimes does not.
BRADY is a career rating. A quarterback's first outstanding season is pulled toward a mediocre past, because from the model's point of view one good year is weak evidence about a player with three bad ones. Awards are for the season alone. These are different questions and the answers should differ.
Shrinkage also pulls extremes toward the middle by design, and awards are given to extremes. A rating built to still be true next season will always give some of this away. That is a statement about what the rating is for, not a defect waiting to be tuned out.
Where a quarterback won an award, it is marked in the tables. Where the rating puts him fourth, that disagreement is shown rather than smoothed over.
What It Deliberately Gives Up
There are two ways to grade a quarterback rating, and they pull against each other.
One asks whether it predicts how much the offence will score. The other asks whether it is describing the quarterback rather than the team around him.
A rating can do very well on the first by quietly measuring the supporting cast — a good offensive line and good receivers really do predict scoring — while getting worse at the thing it claims to measure. So BRADY is checked against a second test built on games where a starter got hurt and a backup took over: the roster is the same, only the passer changed.
Three signals that raised the first score were dropped because they failed the second. The clearest was a measure of how pass-heavy a team's play-calling was. It looked like one of the strongest inputs in the model. It was forecasting the scoreboard: a worse quarterback means a team behind, and a team behind throws more.
Removing it made the rating measurably worse at predicting offensive production, and measurably better at describing a quarterback. That trade is the reason for the whole exercise.
Wins Above Replacement
BRADY says how good a quarterback is per play. It cannot say what he was worth — being excellent for four games is not the same as being excellent for seventeen.
WAR counts the games. For each one he started, it asks how many more wins an otherwise average team would expect against an average opponent with him instead of a replacement-level quarterback. A game he left at half-time counts as half a game.
The conversion from rating to wins is the quarterback's whole contribution to his team. Across every team-game since 2003, a point of BRADY goes with 31 points a game of team strength, and a starter's BRADY accounts for about 40% of how good his team is — close to the third the research on the position arrives at. That relationship has not changed from one era to the next.
Each season is measured against itself: a quarterback is compared with that season's median starter, and a replacement sits a fixed distance below it. So a 2007 season and a 2023 season are judged by how far each stood above the quarterbacks of its own year, and a legendary season from the 2000s reads as one.
Seasons are compared on WAR per game started, so a sixteen-game season stands beside a seventeen-game one, and a season needs eight starts to be ranked. The career total — every game, added up — appears on one page only: the all-time career board, where longevity is the point.
It deliberately does not multiply by how often he threw. A quarterback's busiest games are his losses — teams throw when they are behind — so a WAR that counted dropbacks paid passers for trailing. And a passer who throws fifty times a game moves the result no more than one of the same rating who throws thirty: scaling the rating by volume makes RICE's predictions worse, not better.
Replacement is what quarterbacks who were second or lower on their team's depth chart actually produced in the games they started: 0.127 expected points a play below the season's median starter. The obvious alternative — everyone who threw fewer than 250 times in a season — is contaminated, because it sweeps in every starter who got hurt in September, and those men are good.
The median ranked season is worth about 0.12 wins a start over a replacement, in every era. A very good one is 0.24, and the best on record — Drew Brees's 2011, Tom Brady's 2007 and Peyton Manning's 2004 — are about 0.36: some six wins over seventeen starts. Across a career, Tom Brady is 86 wins above replacement.