Preview · live names, live recommendations, and signal contribution stay Pro. Calibration and the season record are public.
Browse the desk shell before subscribing. Premium values are redacted server-side, not merely blurred in the browser.
2026 current season record
The same settled current-policy aggregate used by the root and Today surfaces. Static calibration evidence is identified separately below.
STRONG is the conviction tier. It has lagged LEAN on price-reference ROI this season. Hit rate is not profit. The delayed CSV is the receipt.
2025 historical benchmark
The historical benchmark stays separate from the current 2026 settlement record above.
2025 is the historical book, not a victory lap. Price-reference ROI finished negative. Hit rate is not profit. The delayed CSV is the receipt.
STRONG versus LEAN is the conviction split, not a marketing grade. STRONG has lagged LEAN on the stored price-reference ROI. Open the delayed ledger if you want the rows.
Live recommendations, player names, and signal contribution stay with Pro. Calibration and the public W–L are not a paywall.
Predicted probability vs actual rate.
The curve is public. Live names are not. This is the audited current-policy seed through 2026-06-06; the W–L above is the live season record.
Predicted probabilities versus realized hit rates on the audited current-policy seed.
Perfect — prediction matches reality
Posterior — public calibration
- Reading the chart.
- When the model says 0.70, the underlying event should occur roughly seventy percent of the time. The closer the curve hugs the diagonal, the more our stated confidence matches the world.
- Why it matters.
- A miscalibrated model can post a healthy hit-rate while sizing bets at the wrong confidence and bleed bankroll silently. Candidate calibrators — isotonic, Beta, BBQ — are evaluated daily against resolved entries; one is promoted only after minimum-sample, holdout, and stability gates clear.
- Hover a point.
- Each bucket displays its predicted probability, the empirical hit-rate, and the sample size — the bigger the n, the more reliable the bucket.