AI Scorecard
Every prediction. Every result. Nothing hidden, nothing deleted.
Lower is better. A random model scores ~0.33; a perfect model scores 0.
Brier score measures calibration, not just accuracy. A model that says "60% home" when homes win 60% of the time scores perfectly on those calls even if accuracy looks flat.
Context: the best football prediction models sit at 55-65% accuracy across a full season. A Brier score below 0.22 is competitive. Anything above 60% on a large sample is strong. High-confidence picks should outperform low-confidence ones. If they don't, the model is miscalibrated.
High-confidence picks should win more often. If they don't, the model needs recalibration.
When Marcus says 70%, does it actually win 70% of the time? Predicted Actual
This is what honest AI looks like. We don't delete the misses.
The record is public and permanent. If you find the accountability useful, support Marcus on Ko-fi.
Support on Ko-fi