The Great Myths of Projective Accuracy

Ed: BaseballHQ.com founder Ron Shandler first penned this piece in 2004, it soon became a BHQ classic. It was updated once in 2009, but was long overdue for another refresh. With the 2026 season and our end-of-year evaluations under way, we pulled it out again and asked Ron to make some minor updates. Happy offseason as the march to 2027 kicks off.

 

Ashley-Perry Statistical Axiom #3: Skill in manipulating numbers is a talent, not evidence of divine guidance.

Ashley-Perry Statistical Axiom #5: The product of an arithmetical computation is the answer to an equation; it is not the solution to a problem.

Merkin’s Maxim: When in doubt, predict that the present trend will continue.

 

The quest for the most accurate baseball forecasting system continues.

I’ve been publishing player projections for over 40 years. During that time, I’ve gained insight into the work of many skilled analysts and forecasting systems. Despite our efforts to predict the future, some constants remain. The core of every system has largely consisted of the same key elements.

  • Players perform based on their past history and trends.
  • Their skills improve and decline with age.
  • Health, expected role, and environment shape outcomes.

These elements keep projections within a realistic range. They stop us from projecting 40 home runs for slap hitters or 40 stolen bases for the slow-footed. However, within this believable range, there is a gap where accuracy dissolves. Even the most consistent and elite power hitter could hit 40, 45, 35, or even 50 home runs in any given year.

Why? While all systems are founded on the same fundamentals, they are limited by the same constraints. We are still trying to project...

  • a bunch of human beings
  • each with their own unique skill sets
  • each traveling along their own unique arc of growth and decline
  • each having different abilities to resist and recover from injury
  • each restricted to opportunities set by other people
  • and each producing a set of statistics heavily influenced by external noise.

Despite intuitively knowing these truths, we still resist them. Baseball is so measurable that even modest predictive success fuels the illusion that a better, more accurate system must be just around the corner. So we build large, intricate models that analyze obscure relationships and try to bring us closer to perfection. However, all this effort only takes us deeper into the abyss.

Why? Because perfection is impossible, and no one seems to have a clear understanding of what success truly is.

 

Measuring success

Is achieving reasonable predictive accuracy even possible? Most analysts agree that only about 65–70 percent of players are “somewhat” predictable in any given year due to injuries, managerial decisions, and other uncontrollable factors. However, even within that 70 percent, there is no consensus on what it truly means to be accurate.

The only completely accurate projection would look like this:

 ABHRRBISBBAOBASLGOPS
PROJ500259515.280.330.450.780
ACT500259515.280.330.450.780

Perfect across the board. Perfectly impossible, given that all these categories move more or less independently for six months. 

If we try to simplify, for example, by evaluating accuracy with a single metric like OPS, new problems emerge. If I project a player’s OPS to be .783, all of the following lines “succeed”:

PlayerABHRRBISBBAOBASLGOPS
Albert55625778.266.327.457.78357
Bill5799556.316.342.440.78260
Charlie37563320.293.357.427.78346
Dave31912492.276.344.439.78262

 To a scientist, these are equivalent. To a fantasy owner, they are not remotely comparable. And OPS ignores playing time, which matters enormously in fantasy contexts.

The Rotisserie dollar metric accounts for playing time and fantasy category weightings, but it introduces its own distortions.

PlayerABRHRRBISBBAR$
Albert55674257780.266$12
Frank5408784680.302$12
Gary539872897110.239$12
Hector57779759330.258$12
Jason50393349010.245$12

Five players can all earn $12 yet deliver completely different category profiles. The last thing your power-rich, speed-starved team needs is for me to project Hector and for you to end up with Jason.  “$12 is $12” is not meaningful guidance.

Instead, perhaps focusing on individual stat categories would provide clearer insight. Not so fast. The question of “What constitutes accuracy?” then becomes personal:

If I project that Jason will hit 45 HRs this year and he only gets 44, you’d probably accept that margin of error. But what if he only hits 43? Or 42? Or 40? Or 39? When do we reach the point where the projection is considered a “failure”?

You might say “40.” I might say, “Okay, so if Jason has 39 HRs on the final day of the season, and he hits a long fly ball that the centerfielder makes an incredible over-the-wall leap to rob him of that 40th homer, has that one event been the difference between success and failure?” We must draw the line somewhere, but there is always a grey area where it can go either way. The size of this grey area varies for everyone.

We tested this at BaseballHQ.com. Here’s what readers said when asked at what point a projection “fails”:

If I projected a player to hit 35 HRs this year, what is the actual HR count at which you would see my projection as failed?

 342% 
 323% 
 3018% 
 2831% 
 2624% 
 2414% 
 225% 
 203% 

 

If I projected a pitcher to win 15 games this year, what is the actual number of wins at which you would consider my projection to have failed?

 144% 
 1310% 
 1233% 
 1127% 
 1017% 
 93% 
 82% 
 73% 

 

 

There is no clear consensus in either poll. Accuracy can only be judged based on your own subjective tolerance for error.

You might say, “There must be some benchmark. There must be some way to gauge accuracy.” But even a universal tolerance threshold — say, 10 percent error — quickly falls apart.

Consider the following projection and outcome:

 ABRHHRRBISBBA
PROJ550791692911313.307
ACT599701692610010.282

 Eyeballing it, this looks close. But every stat category misses the 10 percent threshold. Should we call it a failure? 

What if we use a 15–20 percent tolerance? 

 ABRHHRRBISBBA
PROJ550791692911313.307
ACT6326216923877.267

My personal eyeball test tells me that a 20 percent error is too high for me to accept. However, you might find it perfectly acceptable within your own tolerance for error.

The irony with the above examples is that, despite the shortcomings in batting average, both projections nailed this player’s total hits. 

All of which raises other questions...

  • If a projected slugging percentage is dead on, but the player hits 10 fewer HRs than expected (and perhaps 20 more doubles), is that considered a success or a failure?
  • If a pitcher’s projected hits and walks are accurate, but the bullpen and defense collapse, raising his ERA by a run, is that considered a success or a failure?
  • If a speedster’s predicted success rate for stolen bases is perfect, but his team replaces the manager with someone who doesn’t prioritize running, and the player ends up with half as many SBs as expected, is that considered a success or failure?
  • If a batter is traded to a hitters’ park and all the experts predict an increase in production, but he posts a statistical line exactly as if he had not been traded, is that considered a success or a failure?
  • If a closer’s projected ERA, WHIP, and skills metrics are perfect but he records 20 saves instead of 40 because the GM chose to sign a high-priced free agent, is that considered success or failure?
  • If I project a .272 batting average in 550 AB and the player only hits .249, is that a success or failure? Most will say “failure.” But, wait a minute! The real difference is only two hits per month. That shortfall of 23 points in batting average is because a fielder might have made a spectacular play, or a screaming liner might have been hit right at someone, or a long shot to the outfield might have been held up by the wind... once every 14 games. Does that constitute “failure”?

In the end, accuracy is often less important than strategic usefulness. If your bullpen is loaded, whether your third closer gets 15 or 5 saves might not matter. If you lead the league in HR, your power hitter falling short by five HR probably doesn’t impact much. And a pitcher’s ERA being off by a full run usually affects your total standings points far less than people think.

 

Comparing systems

Assessing accuracy within a single projection set is already tough. Comparing multiple systems becomes even more complex.

Every year, fantasy leaguers rush to find the most accurate forecasting method. “Objective” studies are published to identify the best prognosticators, but these studies are often flawed. Why? It’s nearly impossible to avoid bias in any comparative analysis. Here are a few ways this bias is introduced:

Sample bias: Which sources are included? Is it a comprehensive list or a carefully curated one? Which players qualify? Only those meeting a specific playing time threshold? Is that threshold consistent across all sources? 

Variable bias: Which metrics determine success? OPS? Dollar values? Win Shares? Each option favors different systems.

Methodology bias: Does the study use a recognized, statistically valid method to validate or dismiss variances? Or does it depend on a flawed system like Rotisserie scoring, which distorts the facts by exaggerating small differences and minimizing significant ones?

This is especially true for studies that include the author as one of its subjects (whether openly or through proxy). The reason is simple: a tout won’t publish such an analysis unless they can present themselves in a favorable light. And the only way to do that is to embed some bias into the study's structure.

In the end, objective analysis is almost impossible unless an independent third party designs the study and all forecasters agree to the rules.

Peter “Ask Rotoman” Kreutzer has this view: “Someone who tries to sell you projections that are ‘much better’ than any others is bullshitting you. The important thing for you, as a consumer, is to understand which system your prognosticator uses, what biases it introduces, and how to make the necessary adjustments to incorporate risk evaluation into the process. Only then can you find the players who best fit your league’s rules.”

 

Other challenges to assessing projections

Ashley-Perry Statistical Axiom #4: Like other occult techniques of divination, the statistical method has a private jargon deliberately contrived to obscure its methods from non-practitioners.

Complexity for complexity’s sake: As users of player projections, we want quick, simple answers and are in a hurry to make decisions. We seek trusted sources and rely on them to do the heavy lifting. The more effort we see, the more credible we believe that source is. But this can create false impressions.

One theme I often explore is accepting imprecision in our analyses. This idea might seem counterintuitive, given our increasing knowledge base. But humans are affected by random, external factors; the notion that we can create complex systems to measure these unpredictable elements with precision is what’s truly counterintuitive. 

Research indicates that the simple “Marcel” forecasting system, which averages recent seasons with minor adjustments for age, is nearly as effective as more advanced methods. If 70 percent accuracy is the highest we can realistically expect, Marcel alone achieves about 65 percent. All our advanced systems are competing to capture the remaining 5 percent.

And so, we often obsess over hundredths of a percentage point and treat tiny differences as absolute truths. We declare winners and losers among systems separated by 68 percent versus 67 percent success rates—differences often decided by a few wind-blown home runs and a couple of seeing-eye singles—which are completely irrelevant over a single 23-player fantasy roster. 

We forget simple truths:

  • The difference between a .250 hitter and a .300 hitter is roughly one hit per week.
  • A true .290 hitter can post .254 or .326 in any given year.
  • ERA is heavily affected by sequencing and workload distribution. A pitcher allowing 5 runs in 2 innings will have a different ERA impact than one allowing 8 runs in 5 innings, even though, for all practical purposes, both got rocked. 

Gall’s Law: A complex system that works is invariably found to have evolved from a simple system that works.

Occam’s Razor: When you have two competing theories that make exactly the same predictions, the one that is simpler is preferred.

Systems that try to impress us with their complexity as proof of their credibility might be no better than a room full of monkeys with spreadsheets. At the very least, they produce projections that are ‘close enough’ for our needs and generate draft results that are nearly indistinguishable from those of a simian-driven system.

Married to the model: Beware if a tout is so devoted to his forecasting model that “it” becomes more important than the projections.

Whenever I hear an analyst write, “Well, the model spit out these numbers, but I think it’s being overly optimistic,” I cringe. Well then, change the numbers! The mindset is that you have to cling to the model, for better or for worse, to legitimize it. The only way to change the numbers is to change the model. Is the goal to develop the best model or to generate the most accurate projections?

The Comfort Zone: In October, reality is black or white. In March, it’s all shades of gray. But it’s much easier for fantasy leaguers to draft their teams using blacks and whites, so analysts have to commit. Gray is out, even when a projection is highly uncertain.

For example, if a consistent 40-HR hitter has a poor season, hitting only 25 homers, the obvious questions are, “Was this an anomaly? Will he recover? If so, how much will he recover?”

The typical forecast would never venture into uncharted, sub-25 HR territory, treating the off-year as the start of a trend. Neither would a computer projection see a full rebound to 40-HR levels because the last season couldn’t simply be ignored. A typical projection would rely on history and likely split the difference, perhaps giving recent performance a slight edge.

Most analysts cluster their projections within a narrow range of publicly acceptable outcomes—a comfort zone. They avoid extreme values. Even when evidence suggests an outlier is highly possible— whether positive or negative—they drift toward the middle.

These projections present a cautious view of the future. This is the safest spot if this player repeats his 25-HR performance or rebounds to 40. The projections fall within a comfortable range where neither outcome would be too far off.

I’m not suggesting that living in the comfort zone is a bad idea. However, I’d remind you that fantasy titles are won by those who venture outside of it. 

The Hedge: The hedge is used to avoid taking a firm stance rather than committing to anything, often appearing in a player commentary. In this way, the hedge acknowledges the “grays.” An example:

“He’ll probably bounce back from this career low, but playing home games in a pitchers' park won’t help his hitting stats. Still, he’s just a year removed from a 40 HR season and will be only 30 years old. He might be a bargain.” 

Viewing expectations this way can help you stay open-minded about the wide error margins inherent in the process. But be aware that some use this tactic simply to avoid providing any meaningful analysis. It’s fence-sitting disguised as insight.

The Outliers: In any fantasy league, the winners are often those who include the most outliers on their rosters. A BaseballHQ poll supports this:

In what place was the team that owned (this season’s biggest surprises)? 

 Winner!25% 
 2nd or 3rd24% 
 4th or 5th21% 
 Lower than 5th30% 

One in four owners finished in first place! About half placed no lower than third, and 70 percent finished no lower than fifth! While these results are impressive, performances like these are the hardest to predict. Still, the prognosticators who perform the best in this exercise probably deserve their props, shouldn’t they?

According to analyst John Burnson, the answer is no. He states: “The issue is not the success rate for one player, but the success rate for all players. No system is 100 percent reliable, and in trying to capture the outliers, you weaken the middle, thereby losing more predictive power than you gain. At some level, everyone is an exception!”

Peter Kreutzer again: “Those projections that are outside the comfort zone, as Ron calls it, are flashy but of little statistical use. The predictor who gets the general flow right—guys who improve, guys who fall off—will make you money.”

And that might be the single most meaningful measure of accuracy.

 

Finding relevance

Berkeley’s 17th Law: A great many problems do not have accurate answers, but do have approximate answers, from which sensible decisions can be made.

Our BaseballHQ projection system is built on this philosophy. It begins with objective baselines — five-year trend analyses, minor-league translations, aging curves, and more. Then it provides “skills flags”— moments when a player’s skills or trends don’t align with his surface stats. After that, the process becomes more hands-on and deliberately subjective.

We take these positions:

The best projections are often those that are just enough outside the expected range to influence decision-making. In other words, it doesn’t matter if I project Player X to bat .320 and he only hits .295; what matters is that I projected .320 while everyone else projected .280.

We might consider evaluating projections based on their practical value rather than just accuracy. For instance, is it more meaningful to tell you that a consistent .300 hitter will hit .300 (but actually hits .290), or that a steady .250 hitter will improve to .290 (but hits .270)? By the end of the season, the first projection would have been more accurate, but the second — despite being off by twice as much — could have been more valuable. The second projection could have influenced you to bid an extra dollar on Draft Day, leading to greater profit.

In short:

  • A projection slightly outside consensus can be more valuable than a “safe” number.
  • Intrinsic value matters more than precise accuracy.
  • Insight that influences your draft-day behavior is more important than statistical perfection.

Ultimately, my main goal isn’t strict mathematical accuracy. It’s to affect how you behave on draft day. 

For certain players with notable underlying skills metrics or trends, we often publish projections that aren’t meant to show a “most likely case," but rather a “strong enough case to influence your decision." Sometimes, there are reasons to step outside the comfort zone.

If a player has underlying indicators pointing to meaningful upside, publishing a $27 projection may nudge you to bid $22 when the room stalls at $21. A “safer” $23 projection is less likely to move you. Prices, after all, are not intrinsic—they’re market-driven. If you believe that a player is worth $26 and land him for $21, you'll have overpaid if the rest of the league sees him as no more than a $17 player, even if he's actually worth $35. My goal is to help you understand how to leverage that volatility and identify your real profit opportunities.

In snake drafts, the same logic applies. If ADP pegs a player as a 14th-rounder but we believe his skills merit a 10th-round valuation, you risk missing out on him if we project a 13th-round value. Projecting him as a 10th-rounder gives you the confidence to grab him in round 12 and still profit. Especially in snake drafts, where two-thirds of all picks are losers, there is far less precision than we are led to believe.

And I want you to make these decisions with as little hesitation as possible. That confidence comes from the trust I work to build between us, rooted in sound analysis and a 40-year track record that has proven effective.

All of this answers the question, “For any player, what is the one piece of information more important than the most accurate projection?” That information is how the other owners in your league value that player. If you understand that and have a sense of a player’s potential, it doesn’t matter at all how accurate your projections are.

And with a 70 percent ceiling on accuracy, dollar values and round projections are inherently fuzzy. Two-thirds of players will finish within plus-or-minus $5 of their projection. A 9th-round player might produce anywhere from 6th to 13th-round value. Yes, 70 percent accuracy means a player’s actual value can vary by as much as $10 or seven rounds from the projection. That’s the best case, on average.

We’re not intentionally producing inaccurate projections. We’re simply exploring potential scenarios at the edges of the comfort zone, based on solid underlying indicators, all designed to exploit market behavior. We are publishing useful projections.  

If that makes me a sabermetric hack, so be it. Our readers win their leagues, so I’ll accept the baggage.

Winning is everything. 

While reasonably accurate projections matter, they only take you part of the way. Even perfect projections wouldn’t guarantee victory. Fantasy baseball isn’t played in a vacuum. You must operate within the constraints of league rules, draft dynamics, and other owners' decisions.

Even with a crystal ball and perfect knowledge of each player’s stats next year, you can still lose.

How is that even possible? Accurate projections are just one part of the puzzle. Valuation and roster building are even more crucial for success. We’ve done retro-drafts – where everyone already knows each player’s stats in advance — and excelling in these events requires a specific skill that perfect stats alone can't provide. BaseballHQ.com offers those tools as well!

We believe your goal is to win. This game is about what you know within the context of what everyone else knows. Our system helps you navigate that context more confidently. So, if we project your slugging shortstop to hit 39 HRs and everyone else is projecting 42, remember that the difference between their projection and ours might be three errant gusts of wind.

Baseball Variation of Harvard Law: Under the most rigorously observed conditions of skill, age, environment, statistical rules and other variables, a ballplayer will perform as he damn well pleases.

 

More From Research

The pitching landscape has shifted yet again, and our Pure Quality Start metric undergoes a minor shift to level-set the results.
Nov 28 2023 1:01am
With the new pitch clock we examine who the clock may impact and how the clock may or may not impact pitcher performance.
Feb 23 2023 1:05am
Several years have passed since the original article. It's time for an update and a look at the initial 2021 list.
Jun 17 2021 12:04am
The elite rate at which Dansby Swanson produced 95+ mph exit velocity in 2019 suggests that he's an overlooked breakout target.
Jun 15 2020 1:05am
Nick Pivetta gave up less very hard contact in 2019 than you might think. Could it be a precursor to a rebound season?
Jun 9 2020 1:05am

Tools