TSSE was built in one offseason: eighteen seasons of results, a rating engine, and a research series, all in the weeks since the 25-26 season ended in June. Before the girls kick off in August, that research goes back into the engine. This is the changelog, and a note about what we're keeping under the hood from here on.
The first version of this rating system treated every program the same way twice: at birth and every offseason. A brand-new school started at 1500 plus a small class adjustment, whether it opened in a booming Williamson County suburb or a county that had never fielded a varsity team. And every June, every program's rating slid 25 percent back toward the mean, whether it graduated two players or eleven.
Both of those were reasonable defaults when the engine knew nothing. It knows things now. An offseason of reports established, with numbers, that a program's context predicts where it will settle (The Zip Code Effect), that keeping a core together beats renting one (The Freshman Report), that program experience matters where roster age does not (The Experience Report), and that the coach is the thread that holds the whole thing across graduation cycles (The Coaching Carousel). An engine that ignores its own site's findings is leaving accuracy on the table. So, two changes.
When a new program enters the pool, TSSE now hands it a prior instead of a blank slate. The logic is the same one a good scout would use before ever seeing the team play: what class will they play in, what division, and what does the neighborhood around the school look like?
The evidence for that last part is the strongest single finding on this site. Among public programs in the v1 ratings, the wealthiest quartile of neighborhoods carried an average rating 169 points above the poorest quartile in boys and 188 points in girls, and the private-public gap runs another 80 to 90 points on top of class effects. None of that is destiny. Bearden's quarter-century of coaching continuity beats its zip code, and Tri-Cities beats its own SEI every season. Which is exactly why the prior only takes a fraction of the observed gap:
| Component | Rough weight | Where it comes from |
|---|---|---|
| Base | 1500 | Every program, both pools, unchanged. |
| Class & division | ≈ −50 to +120 | Historical cross-class results; the D2-AA privates enter highest, 1A enters lowest. Carried over from v1. |
| Neighborhood (SEI) | up to ≈ ±45 | About a quarter of the observed quartile gap. Weighted slightly heavier for girls, where the money signal is stronger. |
| Hard cap | ±150 total | No program starts more than 150 from center, ever. |
A prior is a starting guess, not a verdict. New programs keep the accelerated update weight over their first 20 games, so a team that outplays its neighborhood gets credit fast. The point of the change is narrower than it might sound: for the first month of a new program's life, the engine now guesses like someone who has read our research instead of someone who hasn't.
The offseason is where high school ratings go to die. Every senior class walks at graduation, and a flat 25 percent slide toward the mean was our way of admitting we didn't know who lost what. But we track rosters now, season over season, and the rosters say the flat rate is wrong in both directions.
I paired every team-season roster with the rating change that followed it: 6,141 team-season pairs across twelve years. The share of a roster that graduates predicts the next season's rating change at r = −0.24, which for one offseason variable is a loud signal. Split into fifths, it looks like this:
The coach matters too, independently. Programs that changed head coaches between seasons behave like an extra three points of regression compared to programs that kept the same voice on the sideline. Put the two together and the spread is stark: a team returning its core under the same coach holds about 93 percent of its distance from the mean; a team that graduates heavy and changes coaches holds about 87. A flat rate splits the difference and gets both teams wrong.
So the offseason slide is no longer one number. From the 2014-15 season forward, as far back as the roster registry reaches, each program's regression is set from what it actually loses; older offseasons keep the flat rate, since there is no roster to read:
| Term | Rough weight | What it reads |
|---|---|---|
| Base regression | ≈ 18% | The floor case: an ordinary roster, ordinary turnover. |
| Graduation term | ≈ ±8 pts of regression | Scales with senior share relative to the state median (~24%). Heavy senior classes regress more; returning cores regress less. |
| Coach change | ≈ +4 pts | A new head coach between seasons, from the same roster registry behind The Coaching Carousel. |
| Bounds | 10% – 35% | Hard clamp. No roster data for a season means the old flat rate applies. |
One deliberate limit: the regression reads how many players a program loses, not who they are. This site does not track individual minors by name, and the rating system will not either. Team-level counts get us most of the signal with none of the discomfort.
The first cut of this engine barely noticed the state tournament. It read each game's importance from a label in the source data, and that label is unreliable for the postseason, so roughly four in five state tournament games were carrying no extra weight at all, 444 of them tagged as ordinary regular-season games. The matches that decide the season were moving the ratings the least. That is backwards, and it is why the final list read wrong: a champion could land below the team it had just beaten in the final.
The fix has two halves. First, correctness: every postseason game is now identified straight from the official TSSAA bracket registry, matched by the two teams who played rather than a guessable label, so a state final is a state final no matter how the feed tagged it. That alone recovered over a thousand tournament games that had been under-weighted.
Second, the philosophy, and here I took a page from FIFA's book. A goal at the World Cup moves a national team's rating far more than a friendly does, because it is the best teams at the moment that matters most. Tennessee's state cup is that moment. So:
The point was never to force a particular team to #1. It was to let the season's canonical moments carry canonical weight. What fell out is what should: every 2025-26 boys state champion now sits in the top nine overall, each atop or near the top of its class.
Bearden, for its part, did not crater: it slid from #1 to a close #2, exactly the cushion a program with its history should have. And critically, weighting the state cup more heavily did not cost prediction. The backtest is a hair better than before, because those decisive games had been under-counted; counting them properly made the whole ladder sharper. The soft landing was tuned, not guessed; it is the setting that made champions canonical without letting the top of the table run away.
The last change is the quietest and, under the hood, the deepest. Every rating on this site is really a claim about a probability: if Ravenwood is a hundred points above you, how often does Ravenwood win? For years the engine answered that question with the textbook chess number, a 400-point scale, even though the backtest had shown, clearly, that Tennessee soccer is more decisive than chess. The gaps here are wider; a lead means more. We had already fixed the read-out for that, quoting honest probabilities off a steeper curve. But the engine itself was still updating on the shy 400-point math. It was scoring games as if every result were a closer call than it really was.
That one mismatch was the source of a real complaint: the ratings were too generous to a team that just beat up on weaker opponents. Thumping a side you were always going to thump paid more than it should, and holding a heavy favorite to a draw cost that favorite less than it should. So the season-long math rewarded a fat record against a thin schedule. This change closes that gap: the engine now updates on the same steep, real-world scale it reports on. An expected win is worth almost nothing. Beating a genuine peer, or holding a favorite to a tie, is worth a lot.
The effect on the boards is exactly what it should be. A team like South-Doyle, whose 21-2-2 was built beating up the bottom of its region with its best win over a middling side and a loss to Halls, slides back to where the eye test always had it. Unbeaten teams that never tested themselves ease off the elite line. The clubs that actually played, and beat, other good clubs hold their ground. Nobody's ordering was thrown out; the schedule simply stopped being free. And because the engine now scores games at their true difficulty, the backtest is the sharpest it has ever been. This was the last missing piece of v2, the one that makes the number mean what the schedule report says out loud.
Version 1 of this system was documented down to the update formula, and I think that was right for building trust while the site was new. It has a cost, though. A fully published formula invites schedule-gaming (knowing precisely which opponent moves your rating most), it makes every future tuning decision a public re-litigation, and it hands the whole thing to anyone who would rather rebrand it than link it.
So TSSE is adopting the same posture as every serious rating system, from FIFA's SUM formula to Nate Silver's club ratings: the structure, constraints, and rough weights are public (this page is that disclosure), and the exact constants live behind the curtain. You can always check the output against reality; every rating, every game, and every dataset on this site stays free and open.
None of this is a patch stapled onto the end of the list. In July 2026 the engine re-walked every game since 2008 under the new model: each program that ever entered the pool entered with its new prior, every offseason the roster registry can see applied the roster-aware regression, and every postseason game carried its proper state-cup weight. History is re-rated, not rewritten; every result stays exactly as it was played.
The outcome moves like a recalibration, not an earthquake. The new list correlates with v1 at 0.997 in both pools, and the average ranked program moved about 15 points in boys and 22 in girls. The programs that moved most are exactly the ones the research and the trophy case say should: