Methodology

How RICE Works

What RICE Measures

RICE is the Recursive Indicator of Competitive Edge. Recursive, because each week's rating is last week's, updated by what happened since. An indicator of competitive edge, because the gap between two teams' ratings is the point: it should say how likely each is to beat the other, this week and in the weeks that follow.

RICE is a rating of how good an NFL team is right now, on a scale where 75 is the median team-season and the best ever reads 99. It is built to answer one question as well as it can: if this team played an average opponent on a neutral field, what would happen?

The rating itself combines three things — how strong a team has shown itself to be, who is at quarterback, and who is unavailable — with a single fitted weighting. That combination is what gets graded, because on the one question a team rating is judged on, it is the best thing here.

Two further numbers take it apart, both in points per game, both signed so higher is better.

O-RICE is what the offence adds. A team at +5 would score about five points a game more than an average offence would against the same defence. D-RICE is what the defence prevents. A team at +3 would hold an average offence about three points a game below what it usually manages.

They add up to the expected margin — the number of points this team would be favoured by against an average one. A team at +7 is a strong team; the best single season on record is a little over +11.

O-RICE and D-RICE are a decomposition, not the rating. They reconstruct about 97% of it, which is what makes them worth showing; the 3% they miss is split between them so the two columns still add to the third.

Where the Numbers Come From

Three things feed the rating, and every one of them is known before kickoff, so a prediction never contains the game it is predicting. The ratings the tables and charts show are the other side of that: each is a team as it stood after its latest game, with the quarterback who played it.

How strong the team has been. Each team carries one number, the margin it would beat an average team by, and every final score updates it. How far one game moves a team is not a tuned constant: it depends on how much is already known about the team and how much one margin can say. A fumble is charged its average consequence rather than whichever way the ball bounced. Between seasons a team keeps about sixty percent of what it was, because rosters turn over.

Richer measurements were all tried as this signal and lost to the plain margin: expected points per play, points per drive, special-teams value, and every combination of them. They survive in one place, splitting the rating into offence, defence and special teams for O-RICE and D-RICE.

Who is playing quarterback. The starter's BRADY rating: the predictive version, tuned for the games still to come rather than for a career (How BRADY works explains the two). A quarterback in his first dozen starts after taking over mid-season is marked down by a fitted amount, because passers in that spot play measurably below their rating. This is the single largest thing that separates RICE from a rating built only on team results, and it matters most in exactly the games where team results are least informative — the week a starter gets hurt.

Who is unavailable. The snaps a team is losing to players who will not play: everyone on the gameday inactive list, and the players on Friday's injury report, weighted by how often each designation actually means a player sits. This is counted, not estimated, and it only exists from 2016, when snap counts begin.

How a Game Is Predicted

A football game is not a coin weighted by the difference between two ratings. It is about eleven possessions a side, each of which ends in a touchdown, a field goal or nothing.

So that is how it is modelled. Each team's offence is matched against the other's defence to set the rate at which its drives score, and the two resulting score distributions are combined to give the probability of every possible final margin.

How many points a game will have starts there and is then set against where it is played. Under a dome or a closed roof, games score more; in the open, wind takes points away, and a game in fifteen miles an hour or more scores three or four fewer than the same teams would in still air. The weather is the forecast for kickoff, or, for games already played, what it was. Cold on its own matters much less than people expect once the wind is accounted for. None of this changes who is favoured: weather slows both offences alike.

This is why the model has opinions about the shape of a result and not just its direction. Three-point and seven-point margins come out common because the scoring rules make them common, not because anything was fitted to make them so.

January Counts

The playoffs are in everything here: the ratings, the predictions, the tables, and the score the model is built to minimise. They used to be left out, which was wrong. They are the games played by exactly the teams a rating has the most to say about, against a field that is better than average by construction, and they are the games most people weigh most heavily.

They are also harder to call: priced from the end of the regular season, a playoff game scores about 0.642 of log loss, against 0.623 over a team's next eight games, because a field of good teams is closer together. Counting playoff games for more when the ratings update was tested for both RICE and BRADY and did not help, so a January game counts the same as a September one.

What It Is Not

It has never seen a betting line. Point spreads and moneylines are used in one place only — as a benchmark to check the model against — and never as an input. A rating that could see the market would be measuring the market.

It only knows about injuries it has been told about. Until the gameday inactive list is published, ninety minutes before kickoff, the availability component works from the Friday injury report; a game whose team has filed no report is predicted from the other two components alone.

It is worse than the market. Over 5,295 held-out games the closing line is still meaningfully better at picking winners. A model that beat the market would be a surprising claim, and this one does not make it.

The split is a diagnostic, not a better model. Predicting games from O-RICE and D-RICE instead of from the rating itself costs about 0.002 of log loss — small, but large enough to measure, and not in the split's favour.

Honest Performance

Held-out means the model was fitted only on seasons before the one being scored, every time, for every season from 2000 to 2025 — its fixed constants included. The ratings this site publishes are then refitted on every season there is, so a past season's page describes that season with the benefit of everything since, rather than recording what the model forecast at the time. The numbers below are the forecasts.

log lossaccuracy
guessing 50/500.69350.0%
home field only0.68655.9%
RICE0.61865.5%
closing moneyline0.61066.3%

That is next week's game. What the model is actually built for is further out: before each of a team's games, its next few, playoffs included, with each game's quarterback the one who would be named to start it by kickoff. Over a team's next three games RICE's log loss is 0.617, from 0.613 on the first to 0.619 on the third; over the next eight it is 0.623, reaching 0.630 on the eighth, because more happens between now and December than between now and Sunday.

Expected margin correlates 0.40 with the final margin, steady between 0.31 and 0.49 across the seasons. The average miss is about ten points, which is roughly what an NFL game's own randomness costs anyone: the spread of results around any honest prediction is about fourteen points.