What Is WAR in Baseball? How It’s Calculated, What Counts as Good, and the 2025 MLB Leaders
WAR is the number that ends arguments and starts them. It compresses hitting, pitching, fielding, baserunning and position into a single figure — how many wins a player added over a replacement-level call-up — which is exactly why it gets quoted in MVP debates and exactly why it gets misused. Here is what actually goes into WAR, what a good number looks like, where the 2025 leaderboards landed, and how much precision the metric can honestly bear.
Table of Contents
Key Takeaways
- WAR asks one question: if a fresh minor league call-up had taken this player’s place, how many fewer games would the team have won?
- It puts pitchers and hitters on the same scale, which no traditional stat does. A 5.0 WAR season is a star; only 10 to 15 players clear it in a typical year.
- FanGraphs’ fWAR and Baseball-Reference’s bWAR use different engines — FIP versus actual runs allowed for pitchers — so the two never match exactly. Never mix them in one comparison.
- Shohei Ohtani posted 7.5 fWAR as a hitter in 2025, fourth among position players, plus 1.9 on the mound for a 9.4 combined mark — second in MLB behind Aaron Judge’s 10.1.
- WAR is an estimate, not a measurement. Gaps of 0.1 to 0.5 are noise, and single-season defensive WAR is the shakiest input of all.
What WAR Actually Measures
WAR stands for Wins Above Replacement. The replacement player is the hypothetical bench body a team could summon from Triple-A at essentially no cost — not a league-average player, but the guy who fills in when someone gets hurt.
So a 5.0 WAR season means that player was worth about five more wins than that call-up would have been. The unit is wins, which is the point: it converts every kind of contribution into the currency teams actually care about.
That common scale is WAR’s real innovation. Home run totals cannot be compared to win totals, and neither says anything about defense. WAR lets you ask who contributed most this season without first deciding whether to weigh a shortstop against a closer.
It also means WAR sometimes disagrees with the hardware. A player who leads the league in batting average but is a liability in the field grades out lower than a mediocre hitter who is elite with the glove. That gap between reputation and value is the whole reason the metric exists — and the reason people fight about it.
One practical note before the numbers: there is no single WAR. FanGraphs publishes fWAR, Baseball-Reference publishes bWAR (also called rWAR), and the same player in the same season will get different figures from each. Everything below uses fWAR unless stated otherwise.
How WAR Is Calculated: Hitters vs. Pitchers
Position players: bat, legs, glove, and a positional adjustment
For hitters, WAR is built from three contributions.
Offense comes first, measured through wOBA and wRC+ to establish how many more runs the player produced than a league-average bat. Baserunning adds credit for stolen bases and aggressive advancement. Defense is graded with range-and-error metrics such as Ultimate Zone Rating (UZR).
Those three are summed, then adjusted for two things: how hard the position is to play, and how much the player actually played.
The positional adjustment matters more than casual readers expect. Shortstops and catchers are harder to staff than first basemen and designated hitters, so identical offensive lines produce different WAR totals depending on where a player stands. That is scarcity, priced in.
Pitchers: why fWAR and bWAR disagree
FanGraphs builds pitcher fWAR on FIP — Fielding Independent Pitching — which judges a pitcher on strikeouts, walks and home runs alone. Strip out the defense behind him and what remains is closer to the pitcher’s own doing. From there the model estimates runs prevented against an average pitcher, then adjusts for innings and park factors.
Baseball-Reference takes the other road. Its bWAR starts from RA9, the runs the pitcher actually allowed, defense included.
Neither is the correct answer. They encode different beliefs about what a pitcher should be held responsible for, which is why a pitcher backed by a great defense tends to look better in bWAR than in fWAR.
Underneath all of it sits the same comparative question: how did this player do relative to average, and relative to replacement? A .250 hitter with a standout glove can post a strong WAR. A .300 hitter who cannot field will see his sink.
What Counts as a Good WAR? A Season-Long Benchmark
| WAR | What it means |
|---|---|
| Below 0 | Worse than replacement level — actively cost the team |
| 0 to 1.0 | Replacement level: bench piece, roster filler |
| 1.0 to 2.0 | Useful contributor: back-end starter, key bench bat |
| 2.0 to 3.0 | Average everyday regular |
| 3.0 to 5.0 | Quality regular, All-Star candidate |
| 5.0 to 6.0 | Star, inside the MVP conversation |
| 6.0 and up | Elite, legitimately MVP-worthy |
Across a 162-game MLB season, only about 10 to 15 players reach 5.0 WAR. Clearing 2.0 already means a player contributed above the league average.
Cross 3.0 and a player has performed at an All-Star level. Cross 5.0 and he belongs in MVP discussion. Cross 7.0 and he single-handedly moved his team about seven wins in the standings — the tier where MVP stops being an argument.
The truly historic seasons go past 10.0. Barry Bonds hit 73 home runs in 2001 and posted 12.5 WAR as a hitter, one of the highest marks the sport has recorded.
WAR can also go negative. Poor defense or limited playing time can drag a player below zero, which reads as a blunt verdict: the roster would have been better off with someone else in that spot.
And WAR frequently diverges from actual MVP voting. Ballots reward visible counting stats — home runs, RBIs — while WAR folds in defense and on-base skills. When the two disagree, the disagreement is usually about which side of the game you think deserves more weight.
NPB Has WAR Too, But You Can’t Compare It to MLB
Nippon Professional Baseball tracks the same concept, and it works the same way internally. What does not travel is the comparison.
WAR is defined against a league’s own average and its own replacement level. NPB plays a 143-game schedule at a different competitive level, so an NPB WAR figure and an MLB WAR figure are measured against different baselines.
This is why a player’s WAR shifts when he moves from NPB to MLB even if his underlying skill has not changed. The player stayed the same; the yardstick moved.
Shohei Ohtani’s 2025 WAR: Two-Way Value Breaks the Scale
Ohtani’s peculiarity in WAR terms is structural: he is the only player stacking pitching WAR and hitting WAR in the same season.
As a hitter in 2025, the Dodgers star posted 7.5 fWAR — fourth among all MLB position players. Aaron Judge led at 10.1, followed by Cal Raleigh at 9.1 and Bobby Witt Jr. at 7.7.
He was not a full-time hitter, though. Ohtani returned to the mound during 2025 and added 1.9 pitching WAR on top of the bat, bringing his combined total to 9.4 — second in MLB across all players, behind only Judge.
That total matches the 9.4 he put up with the Los Angeles Angels in 2022 — and the 2022 version got there the hard way, throwing 166 innings on the mound while appearing in 157 games as a hitter. A two-way season running at full throttle, with no precedent in the sport’s history.
Landing on 9.4 again in 2025 says the ceiling has held. Since the two-way act became a full-time proposition, Ohtani has stayed in MLB’s top tier.
The structural point is worth sitting with. Every other player can accumulate value in one column. Ohtani accumulates in two, which means he is competing for WAR in twice as many places as anyone he is being compared to.
2025 MLB WAR Leaders (fWAR)
Final 2025 fWAR leaderboards, split by pitchers and position players.
Pitching WAR leaders
| Rank | Player | Team | Pitching WAR |
|---|---|---|---|
| 1 | Tarik Skubal | Detroit Tigers | 6.6 |
| 2 | Paul Skenes | Pittsburgh Pirates | 6.5 |
| 3 | Christopher Sánchez | Philadelphia Phillies | 6.4 |
| 4 | Garrett Crochet | Boston Red Sox | 5.8 |
| 5 | Logan Webb | San Francisco Giants | 5.5 |
| 6 | Jesús Luzardo | Philadelphia Phillies | 5.3 |
| 7 | Yoshinobu Yamamoto | Los Angeles Dodgers | 5.0 |
| 8 | Max Fried | New York Yankees | 4.8 |
Source: FanGraphs WAR Leaderboards (2025)
Skubal topped the pitching side at 6.6. The Tigers left-hander followed his 2024 AL Cy Young Award with 195.1 innings in 2025 and won the award again. His profile — heavy strikeouts, minimal walks — is exactly what FIP-based WAR rewards most.
Skenes finished second at 6.5. Pittsburgh’s young ace threw 187.2 innings and claimed the NL Cy Young Award, building on a rookie season that already looked like an outlier.
Sánchez logged 202.0 innings for 6.4 WAR, establishing himself inside a deep Phillies rotation.
Yamamoto was the top Japanese pitcher at 5.0, seventh overall. The Dodgers right-hander, who signed out of NPB’s Orix Buffaloes (Pacific League, based in Osaka), covered 173.2 innings.
Position player WAR leaders
| Rank | Player | Team | Position Player WAR |
|---|---|---|---|
| 1 | Aaron Judge | New York Yankees | 10.1 |
| 2 | Cal Raleigh | Seattle Mariners | 9.1 |
| 3 | Bobby Witt Jr. | Kansas City Royals | 7.7 |
| 4 | Shohei Ohtani | Los Angeles Dodgers | 7.5 |
| 5 | Geraldo Perdomo | Arizona Diamondbacks | 7.1 |
| 6 | Trea Turner | Philadelphia Phillies | 6.7 |
| 7 | Corbin Carroll | Arizona Diamondbacks | 6.5 |
| T-8 | José Ramírez | Cleveland Guardians | 6.3 |
| T-8 | Francisco Lindor | New York Mets | 6.3 |
Source: FanGraphs WAR Leaderboards (2025)
Judge led both the position player list and the all-player combined list. He turned 663 plate appearances into 10.1 WAR, a total that clears MVP-worthy and lands in historically rare territory — MLB produces a 10-win season only every few years.
Raleigh’s 9.1 in 705 plate appearances is arguably the more remarkable line, because he did it as a catcher. Catchers almost never reach that number; doing so required elite production on both sides of the ball.
Read the two lists together and a pattern emerges. The players at the top are rarely bat-only. They run, they field, they play demanding positions. That combined value is precisely what batting average and home run totals cannot see, and precisely what WAR was designed to expose.
The Limits of WAR: Three Things the Number Can’t Do
Everything above assumes WAR is trustworthy. It is useful, but it is an estimate rather than a measurement, and three limits are worth holding onto.
First, there is no single correct WAR. FanGraphs’ fWAR, Baseball-Reference’s bWAR and Baseball Prospectus’ WARP each choose different defensive inputs and different baselines, so the same player in the same season gets different numbers. Baseball-Reference says so plainly in its own documentation: the calculation runs through hundreds of steps, and dozens of them are places where reasonable people land differently. WAR is closer to a framework than a formula, and it was never going to resolve to one exact decimal.
Second, defensive measurement wobbles. The fielding component leans on UZR and DRS, and those metrics swing year to year across a single season’s worth of playing time — they need multiple years to settle. When defensive WAR moves, total WAR inherits the error. For a single season, especially for a player whose case rests heavily on his glove, read the number as a range.
Third, replacement level is a convention, not a law of nature. In 2013 FanGraphs and Baseball-Reference agreed to align their baselines at a .294 winning percentage, or 1,000 WAR per 2,430 games. That agreement made the two systems comparable, which is genuinely valuable — but it was a decision, not a discovery.
Taken together: WAR is a well-built approximation of a player’s value squeezed into one number. Treat it as an estimate with error bars, not as settled truth.
How to Read WAR Without Misusing It
Don’t crown a winner over a 0.2 gap
If WAR is an estimate, differences of roughly 0.1 to 0.5 sit inside the noise. Declaring one player better because his WAR is 0.2 higher asks the metric for precision it does not have.
FanGraphs itself recommends reading WAR in bands rather than strict rank order. The benchmark table above works the same way — 2.0 to 3.0 for an average regular, 5.0-plus for MVP candidates — as tiers, not hard cutoffs.
Never mix fWAR and bWAR
The two versions are built on different premises, most sharply for pitchers: FanGraphs on FIP inputs, Baseball-Reference on runs actually allowed. Putting one player’s fWAR next to another player’s bWAR is not a comparison. Pick a version and stay in it.
Be careful comparing across eras and leagues
WAR depends on the league average and the run environment it is measured against. NPB and MLB figures cannot be set side by side, and neither can a high-offense era and a low-offense one without accounting for the difference.
Ignore defensive WAR in small in-season samples
Because fielding metrics stabilize slowly, defensive WAR — and therefore total WAR — is least reliable a few dozen games into a season. Treat April and May WAR leaderboards as curiosities.
WAR describes value; it doesn’t capture everything
WAR is a context-neutral account of the value a player accumulated. It does not price clutch performance or leadership, and depending on the version it may miss catcher framing and game-calling.
Use WAR as the opening move in an argument, paired with traditional numbers and scouting judgment. That is closer to how its own builders intend it to be used.
Where WAR Came From
Knowing why WAR was invented makes both its ambition and its ceiling easier to accept.
The founding complaint was that traditional stats — average, home runs, RBIs, wins — could not account for defense, baserunning, positional difficulty, ballparks or eras, which made cross-position comparison impossible. Fans wanted one yardstick that could measure a pitcher against a hitter and a shortstop against a first baseman. WAR is the answer to that question.
The lineage runs back to Bill James, sabermetrics’ founding figure, who tried to express total player value as one number in Win Shares (2002). Keith Woolner then developed Value Over Replacement Player (VORP), adopted by Baseball Prospectus, which measured a player against others at his position and pushed the idea of replacement level into circulation.
WAR took its modern shape around 2008. That November, Sean Smith published research on replacement level and WAR at The Hardball Times, and Tom Tango released the methodology that still underpins most WAR implementations. FanGraphs introduced “win values” in late 2008 — the name WAR came later — and Baseball-Reference added its own version in 2010 through the work of Sean Forman and Smith. The bWAR nickname “rWAR” comes from Smith’s handle, Rally. Tango contributed the win and run expectancy work; Smith supplied TotalZone for defense and the early WAR framework. The 2013 baseline agreement finished the job.
WAR is now embedded in the sport, and many of the writers who built it — James, Smith, Woolner — work in front offices today. The arguing never stopped, though: James himself has publicly criticized parts of WAR.
That is the honest picture of the metric. Not a finished answer, but a yardstick still being revised in the middle of an ongoing fight about what it should measure.