PredictPulse
Support Marcus
MENU

How Marcus makes each call

A fixed-weight blend of market odds and statistical modelling. No adaptive weights, no black box.

The pipeline in one line

De-vigged market odds (75%) blended with a Dixon-Coles stat model (25%), fixed weights, no daily LLM call.

The two signals

75%
Market signal
De-vigged bookmaker odds

Bookmaker odds encode the collective information of thousands of sophisticated bettors. After removing the vig (the bookmaker's margin, typically 3-6%), the remaining probabilities are a strong, information-dense signal. The market reprices injuries, suspensions, and team news faster than any model. This is empirically the best single signal for football prediction, and why it carries the larger, fixed share of the blend.

25%
Statistical model
Dixon-Coles bivariate Poisson

The Dixon-Coles model fits per-team attack and defence strengths by maximum likelihood over 5 seasons of EPL results, with exponential time decay (roughly 1-year half-life) so recent form is weighted more heavily than results from three seasons ago. It includes a low-score correction factor (rho) that captures the draw inflation inherent in football.

The model outputs a full probability distribution over scorelines. From that distribution, it reads off: P(home win), P(draw), P(away win), expected goals for each team, P(over 2.5 goals), P(both teams score), and the three most likely exact scorelines.

How the signals are blended

The two probability vectors are combined as a weighted average (75% market, 25% stat) and renormalized to sum to 1. The weights are fixed, not adaptive. An earlier version tested weights that re-tuned themselves from resolved results, but the data used to tune that re-weighting was fit on World Cup international matches, not club football. Applying it to the Premier League would have quietly miscalibrated the blend, so the weights were locked instead. Match-level accuracy and Brier score are still tracked openly on the scorecard, feeding a future club-specific refit rather than adjusting today's predictions in the background.

blended = 0.75 × Market + 0.25 × Stat
then: argmax → predicted_outcome

Result selection and the draw problem

The predicted outcome is the argmax of the blended probability vector: whichever of home, draw, or away has the highest probability. Draws are rarely the single largest bucket, so the model structurally under-predicts draws. This is inherent to football: the draw probability almost never exceeds the home or away win probability, and lowering the decision threshold hurts accuracy on balance.

The lever for draw accuracy is calibration, not the decision rule. If the draw probability is well-calibrated, it is communicated accurately as a probability; the model just will not pick "draw" as the outcome unless it is genuinely the most likely result.

Why there is no daily LLM call

The World Cup version of PredictPulse used a large language model as a third signal. Ablation testing on 104 resolved WC matches showed that the 2-signal blend (market + Dixon-Coles) had a Brier score within 0.005 of the 3-signal blend including the LLM. The LLM's information advantage was minimal because bookmaker odds already reprice for news faster than a daily LLM run.

The LLM now writes the weekly featured match analysis only. This is a better use of the signal: generating context-rich written analysis for marquee matches rather than adding marginal prediction accuracy at daily compute cost.

Known limitations

  • Draws are structurally under-predicted. This is a property of football, not a fixable bug.
  • The model does not have access to starting lineup information until the official announcement, typically 1 hour before kickoff.
  • UCL group-stage matches for non-EPL sides use Elo-based ratings as a fallback when Dixon-Coles coverage is thin.
  • Confidence scores have historically shown weak discrimination in small samples. The scorecard tracks per-band accuracy to catch this.
  • Very short-priced favourites (implied odds >85%) are likely to win but provide little value in probability terms: the model has less room to add information.
View the scorecardToday's predictions