How the MLB Model Works
A transparent look at the projections, simulation engine, and adjustments behind our MLB predictions. No hype — just the methodology.
What This Actually Is
Summary:
Predictium's MLB model is a full-game simulation engine, not a regression on team stats. Every game is simulated 2,000–5,000 times at the individual at-bat level: each plate appearance is played out in sequence through both batting orders, tracking runners, outs, score, inning, pitch counts, and times-through-order effects. Pitcher removal is modeled (pitch count, performance hooks, TTO penalty), and bullpens are sequenced with fatigue adjustments. Each simulation produces a complete stat line for every player plus the full game outcome — so the model outputs probability distributions for every market, not single point estimates.
Why simulation?
Baseball is a sequence of discrete matchups — a specific pitcher against a specific batter, in a specific park, with specific runners on base. Team-level averages wash that structure out. Simulating at the at-bat level preserves it: platoon advantages, lineup construction, a thin bullpen after yesterday's 12-inning game, and the wind blowing out to center all flow naturally into the final win probability, run line, total, and first-5-innings numbers.
Architecture: Five Layers
From player projections to market probabilities
Layer 1: Bayesian Hierarchical Player Projection Engine
Marcel 5/4/3 baseline + Steamer/ZiPS blend + Statcast enhancement
Output: full posterior distributions per player parameter
Layer 2: Matchup Probability Engine
Bayesian Log5 / Odds Ratio (pitcher x batter x league baseline)
Modifiers: platoon splits, park factors, umpire tendencies, weather
Layer 3: At-Bat-Level Monte Carlo Simulation Engine (core)
Simulate each PA in sequence; track full game state
Pitcher removal, bullpen sequencing with fatigue
2,000-5,000 simulations per game
Layer 4: Distribution Extraction
Aggregate sims into probability distributions for every market:
ML, run line, totals, F5, Ks, hits, HRs, TB, RBI, SB, outs recorded
Layer 5: Live-Betting Extension (future)
Re-run sims conditioned on current game stateDesign Principle
The simulation produces distributions once; everything the site shows — win probability, run line cover rate, over/under at any line, F5 splits — is read off the same set of simulated games. There's no separate model per market that can disagree with itself.
Player Projections: Marcel + Bayesian
What each player's true talent looks like today
1. Marcel Baseline (5/4/3 Weighting)
The foundation is a Marcel-style projection: the last three seasons weighted 5/4/3 (most recent heaviest), regressed toward league average based on sample size. It's famously hard to beat — so instead of replacing it, we enhance it.
2. Bayesian Hierarchical Pooling
On top of Marcel, a Bayesian layer blends external projection systems (Steamer, ZiPS) and Statcast signals — barrel rate, swinging-strike rate, bat speed, Stuff+ — using beta-binomial hierarchical pooling with position-specific priors. The output isn't a single number per rate; it's a posterior distribution that admits uncertainty for small samples. A rookie with 40 PAs gets a wide distribution pulled toward his position's prior; a 10-year veteran gets a tight one.
| Stat | Marcel MAE | Bayesian MAE | Improvement |
|---|---|---|---|
| K rate | 0.0295 | 0.0300 | -1.8% (Marcel already strong) |
| BB rate | 0.0183 | 0.0161 | +11.8% (blend clearly helps) |
| HR/FB | 0.0108 | 0.0113 | -4.9% (high-variance stat) |
Benchmarked on 2025 pre-season projections vs. actuals (n=1,021–1,202 players). We publish where the Bayesian layer helps and where it doesn't — the blend weights are tuned per stat accordingly.
3. Matchup Probabilities (Log5)
Each plate appearance in the simulation needs outcome probabilities for this pitcher against this batter: K, BB, single, double, triple, HR, and out types. We use the Bayesian Log5 / odds-ratio method to combine pitcher rates, batter rates, and the league baseline, then apply matchup-specific modifiers.
Park, Weather & Umpire Adjustments
Context the simulation sees for every game
- Park factors — per batted-ball type and batter handedness, updated annually. Fenway inflates doubles for right-handed pull hitters differently than it inflates home runs; the simulation applies factors at the outcome level, not as a flat team multiplier.
- Weather — temperature, wind speed and direction (out to center plays very differently than a cross-wind), humidity, and precipitation risk, refreshed on a 5-minute loop before first pitch.
- Umpires — home plate umpire zone size and historical K/BB rate impact shift every at-bat's strikeout and walk probabilities once the assignment is known.
- Platoon splits — every pitcher-batter matchup applies L/R splits from the projection posteriors, so a lefty-heavy lineup against a tough LHP is priced in automatically.
- Rest & bullpen fatigue — relievers carry pitch counts from the last three days; a fatigued or unavailable arm changes late-inning run distributions.
Why Distributions Beat Point Estimates
The output format is the product
Most services publish one number per game. We keep the whole distribution, because the right statistical family matters:
- Beta-binomial for counting props (strikeouts, hits, total bases) — baseball counting stats are underdispersed relative to Poisson, and using the wrong family systematically misprices overs.
- Negative binomial for team run scoring — run variance is roughly twice the mean, and the fat right tail is where blowouts (and overs) live.
- Empirical distributions from the simulation for correlated outcomes like RBI, which depend on who's on base when you hit.
What this buys you
The probability of the over at 8.5 and 9.0 and 9.5, the chance the favorite covers -1.5, the odds the game is tied after five — all from the same simulated sample, all internally consistent.
Validation: How We Know It Works
The only test that matters
Every component is validated walk-forward: projections for a given date are built only from data available before that date, and the simulation is scored against what actually happened. Projection-layer accuracy is benchmarked and published above; full market-level backtests (moneyline, run line, totals, F5, props) are publishing on the backtest page as the historical sweep completes.
About These Numbers
Like our NBA model, we only publish walk-forward numbers — no in-sample results, no cherry-picked windows. If a number looks modest, that's because it's real.
Limitations: What This Model Is Not
Honest assessment
It's Not a Crystal Ball
Baseball is the highest-variance major sport. The best team in the league loses 60 games a year. This is a tool for systematic, probabilistic analysis over a long season — not for calling tonight's winner with certainty.
Small Samples Are Hard
Rookies, players returning from injury, and September call-ups have genuinely uncertain talent levels. The Bayesian posteriors admit that uncertainty rather than hiding it — which sometimes means the model is honestly unsure.
Lineups Move Late
Predictions before lineup confirmation use projected batting orders. A surprise day off for a star can move a game's numbers meaningfully — which is why predictions refresh on a 5-minute loop and carry a lineup status flag.
No fluff. No promises. Just the math.
See It in Action
Win probabilities, projected runs, run lines, totals, and F5 forecasts for every MLB game.
Predictium provides analytical tools for informational purposes. All betting carries risk. Gamble responsibly.