Tennessee Soccer Stats
Tennessee Soccer Stats | Validation · Part II | July 2026

Calling the Whole Season

Part one asked whether the ratings were any good. This one goes further: I predicted the scoreline of every 2025-26 game before it was played, boys and girls, then graded myself. The point isn't the grade. It's finding out what actually moves a prediction, so I can fix the weights.

The rules are strict, because they have to be. Every prediction sees only what was knowable before kickoff: the ratings as they stood that morning, each team's record so far that season, the history between the two clubs, who was home. The weights that combine those into a scoreline were fit on seasons that ended before 2025-26 ever started. Nothing peeks. Then I let it call all games and compared every one to what happened. Same process for both sides, all the way through.

The Grade
The current site formula (ELO only) against the full model (ELO plus home field, form, head-to-head, roster and coach), on the 2025-26 holdout. Lower log-loss and Brier are better; those are the honest measures. Accuracy is the eye-test one.

The Three Trials
The weight-picking ladder for 2026-27. Trial one predicts from the rating alone. Trial two adds the sharpeners: home field, head-to-head, and the current streak. Trial three throws in everything else we track: season form, goal-difference form, roster age, senior share, coach tenure.

Were the Odds Honest?
The most important chart in any prediction system. When the model said a team had a 70% chance, did that team win about 70% of the time? On the line means yes.
Calibration: Predicted vs Actual Win Rate
Each dot is a 10-point probability bucket · the diagonal is perfect
BoysGirlsPerfect
The Predictions, Game by Game
A sample of the marquee 2025-26 games, the ones between two strong teams, with the state tournament first. The gold score is what the model called before kickoff; the white score is what happened.

The Autopsy: What Actually Helps
Starting from ELO-plus-home, I added each extra ingredient on its own and measured whether the season's predictions got sharper. Bars to the right of center helped; bars to the left hurt. This is the whole reason for the exercise.
Predictive Lift of Each Feature, Beyond ELO
Log-loss improvement on the 2025-26 holdout (× 1000) · gold boys, green girls

What We're Changing
The conclusion, and the actual point of Part II: the updates going into the win-prediction weights for the season ahead.
Where This Goes Next
Read First

The Validation Report

Part one: whether the current engine beats the old one, and by how much, across every window we can score.

See It Live

Who is the Best Right Now?

The ratings this model runs on, as they stand today. The 26-27 predictions start from here.

Related

The Great Rivalries

Head-to-head history was the one thing that added real signal. These are the series where it matters most.

Method & Sources
  1. Scorelines use an independent-Poisson goals model with a Dixon-Coles correction for low-scoring dependence, fit by maximum likelihood. Win, draw and loss probabilities come from the resulting joint distribution. Exact engine constants stay behind the curtain, as always; what's published here is the evaluation, not the internals.
  2. "The grade" reports accuracy (share of games where the most likely outcome happened), log-loss and Brier score (probability quality, lower better), and for scorelines: exact-score rate, mean absolute error on goal margin and total, and sign accuracy (correct winner among decisive games).
Tennessee Soccer Stats (TSSE) is a personal, independent project and is not affiliated with the TSSAA or any school. A season is a small sample and any single prediction can be wildly wrong; the value is in the aggregate, and in what the aggregate says to change.