Strikeouts methodology
How the Strikeouts model works
Per-start strikeout projection for each starting pitcher, with Over / Under / Neutral lean based on matchup, opponent, umpire, and price.
Inputs
- • Pitcher rolling K%, K/9, FIP, pitch arsenal
- • Opposing lineup K% rolling windows + handedness mix
- • Umpire zone size factor + K factor
- • Park K factor
- • Weather (temp, wind affecting fly balls)
Outputs
- • Projected K count
- • Over / Under / Neutral lean
- • Public-pick or research-only status
- • When available: the sportsbook strikeout line and price the model compared against
Approach
The strikeouts model estimates how many Ks each starter is likely to record from pitcher form, pitch mix, opposing lineup swing-and-miss profile, park, umpire zone, and weather.
Early tracking showed the model was a little too conservative on pitcher strikeouts, so the displayed projection includes a small correction until the next larger retrain absorbs that behavior.
Training data
Multi-year start-level data joined to opposing-lineup strikeout tendencies and recent pitcher form.
New or low-information pitchers are handled more cautiously. The model may still show a projection, but it will not be treated the same as a starter with a meaningful history.
Calibration
Strikeout performance is reviewed separately for Overs and Unders because the model has produced many more Under leans this season.
Until the sample is deeper, the model is intentionally conservative about which K reads appear as public picks.
How a pick becomes official
For a strikeout read to count publicly, the starter needs enough usable history and the model needs to compare its projection to a real sportsbook strikeout line.
If the pitcher sample is too thin, the line is missing, or the matchup information is incomplete, the read stays informational instead of being graded as an official pick.
Known regressions and recoveries
April audit: we tightened handling for new pitchers so low-information starters do not get over-promoted.
April audit: exact-line pushes are now handled as pushes, not wins or losses.
Known biases and limits
- ! Heavily directional — the model produces far more Under leans than Over leans this season. We are tracking whether this is a real market inefficiency, a feature gap, or a structural bias (Roadmap E).
- ! New / debut pitchers without enough history are kept out of the official public-pick group.
Not for use as
- × Predicting one exact at-bat or one exact inning. These are probability models, not play-by-play fortune tellers.
- × Ignoring the sportsbook price. A model edge only matters if the available price is good enough.
- × Replacing late-breaking baseball judgment. Lineups, weather, scratches, and pitching changes can move after the page refreshes.