---
title: "Against Other Metrics — Held Out"
description: "How RICE and BRADY compare with EPA per play, DVOA, Elo and the quarterback ratings you may have seen."
url: "https://heldout.org/methodology/comparisons"
---

Methodology 

# Against Other Metrics

## How These Compare to Metrics You Already Know

### Against EPA per Play

Expected points added is the raw material, not a competitor. It values a single play by how much it changed the team's expected points, and both RICE and BRADY are built out of it.

What a rating adds on top is everything that makes a season's worth of EPA mean something: adjusting for the defences faced, weighting recent games more than old ones, and shrinking the parts that are mostly luck. A team's raw EPA per play tells you what happened. A rating is an estimate of what will happen.

On held-out games, RICE's team-strength component alone reaches 0.630 of log loss and RICE reaches 0.618. The gap between any two serious team ratings is small — which is the honest headline, and the reason every number on this site is quoted with the games it was measured on.

### Against DVOA

Football Outsiders' DVOA compares each play to a league-average baseline for the same down and distance, then adjusts for opponent. It is the closest relative to how O-RICE and D-RICE are built.

Most of what DVOA reports doing was tested directly against this scoreboard. Some of it survived: counting a fumble the same whether or not it was recovered, and carrying special teams as a third unit. Some did not: its success-rate baseline lost to plain expected points, and discounting garbage time made things worse.

### Against Elo Ratings

An Elo rating moves a team up when it wins and down when it loses, by an amount that depends on the margin and the opponent. It was one of the components here until the team-strength component learned to treat one game as one piece of evidence; since then Elo adds nothing to it. On its own it is still a genuinely strong baseline — FiveThirtyEight's published settings reach about 0.63, within a hair of far more complicated things.

Its weakness is that it cannot see a roster. An Elo rating of a team whose starting quarterback was carried off last Sunday is a rating of the team that played with him. That is the gap the quarterback component fills, and it is worth twice as much in games where the starter changed as across games in general.

### Against Quarterback Ratings You May Have Seen

**Passer rating** is from 1973 and uses completions, yards, touchdowns and interceptions. It does not know about sacks, scrambles, or where on the field anything happened.

**QBR** is ESPN's, is opponent-adjusted, and splits credit for yards after the catch. It is a good measure of what a quarterback did. BRADY is trying to do something slightly different — say how good he is, in a way that should still be true next season — which is why it shrinks toward average and will always look more conservative.

**EPA per dropback** is the honest simple baseline and is genuinely hard to beat. On the 8,981 starts where ESPN's QBR also exists, predicting the offence a quarterback's team produces in his next game, the BRADY inside RICE correlates 0.286 with it, EPA per dropback 0.268 and QBR 0.249. Adding QBR to BRADY adds nothing.

### What All of These Share

Every rating on this page is fighting the same problem: a football season is seventeen games, and seventeen games is not very much data about anything. Most of the difference between two good ratings is how carefully they decide what to ignore.

Every figure here is held out — fitted on earlier seasons, scored on games the model had not seen. The full record, failures included, is in `docs/`.

Last Updated: Sep 26, 2026, 3:56 PM PDT © 2026, Aditya Kishore. All Rights Reserved. Club marks are the clubs’ own trademarks. Past marks courtesy of [SportsLogoHistory.com](https://sportslogohistory.com/nfl-primary-logo/) and data courtesy of [nflverse](https://nflverse.nflverse.com/).
