Two kinds of evidence, kept apart on purpose: the LIVE ledger below grades published picks against real results as they settle; everything under it is the out-of-sample record each model was promoted with.
Every published pick is recorded before its event and graded after it. Nothing is deleted, nothing is regraded.
| Sport | Record | Hit | Brier | CLV |
|---|---|---|---|---|
| MLB | 1094–1100 | 50% | 0.231 | +1.0 |
| NFL | 16–16 | 50% | 0.268 | — |
| UFC | 9–8 | 53% | 0.225 | -0.6 |
82 picks recorded and awaiting settlement. Brier: 0.25 is a coin flip, lower is sharper. CLV: probability points gained on the closing line — positive means the market moved toward our number.
A market earns the right to claim an edge by beating closing prices over at least 40 measured picks with a confidence interval clear of zero — thresholds set in code before any data existed. Verdicts advise; nothing promotes itself silently.
! marks a market where the closing line and the money disagree. Both columns are measured; when they conflict, trust the return and treat the verdict as suspect — the “closing” price is the last one archived before the start, a median 3.07 hours early, and a line taken that far out is easy to beat and says little.
| Market | Verdict | Record | CLV | Return | n |
|---|---|---|---|---|---|
| Runs batted in · MLB | edge rejected | 3052–3026 | +1.24pp | -8.0% | 487 |
| Runs scored · MLB | edge rejected | 2931–2680 | +0.81pp |
Higher is better. Muted bars did not reach significance at p < 0.05 on a paired block bootstrap over weeks.
| Market | Learner | CRPS | Baseline | Delta | p | Log loss | Calib. error | Cov 80 | Rows |
|---|---|---|---|---|---|---|---|---|---|
Anytime TD anytime_td | lgbm | 0.167 | 0.183 | +9.0% | < 0.0005 | 0.4997 | 0.0165 | 0.984 | 28,938 |
dst sacks dst_sacks | lgbm | 0.958 | 1.019 | +6.0% | < 0.0005 | 0.6354 | 0.0137 | 0.899 | 3,230 |
dst takeaways dst_takeaways |
Every recommendation is ranked and gated by an estimated edge. These are the graded picks grouped by the edge we claimed at the time, against what a flat unit on each actually returned. If the estimate carries information, the bottom row beats the top one.
| Edge we claimed | Picks | Hit | Return | 95% interval |
|---|---|---|---|---|
| under 3% | 211 | 47.9% | -10.7% | -25 to +4 |
| 3–5% | 706 | 50.7% | -5.4% | -13 to +2 |
| 5–8% | 908 | 51.9% | -4.3% | -11 to +2 |
| 8–10% | 418 | 45.2% | -15.4% | -25 to -6 |
No relationship worth naming. A bigger claimed edge did not return more than a smaller one, on the sample so far. No threshold has been changed on the strength of this table. The bands were fixed before the numbers were read and are not re-cut; a rule moved to fit what has already happened is not a rule.
Every curve on this page comes from the walk-forward harness and describes a model. This describes the published numbers — scored against the append-only record of what was issued and how it settled.
claimed 53.4% · hit 49.8% [49.1–50.5] n=18,831
Over 18,831 settled picks over 39 event days, these claimed 53.4% and hit 49.8%. The claim sits OUTSIDE the [49.1%, 50.5%] interval on the outcome, so the 3.5 point overstatement is more than sampling noise. Read a published percentage with that in mind.
claimed 56.3% · hit 49.9% [47.8–52.0] n=2,243
Over 2,243 settled picks over 35 event days, these claimed 56.3% and hit 49.9%. The claim sits OUTSIDE the [47.8%, 52.0%] interval on the outcome, so the 6.4 point overstatement is more than sampling noise. Read a published percentage with that in mind.
Nothing here is corrected for. Correcting means fitting, and fitting on a record this shallow means fitting to a handful of slates — the record spans 39 event days, and the floor set for touching a published probability is ten. The gap is shown instead.
The curves behind the scalars, from out-of-sample walk-forward predictions — never re-scored in-sample, which would draw flattery. Reliability: claimed probability against observed frequency; a calibrated model hugs the diagonal, and point size is sample count, because a bin of nine and a bin of nine thousand are different kinds of evidence. PIT: twenty bars that sit flat at the dashed line when the whole predicted distribution is shaped right — a U is overconfidence, a hump is timidity, a slope is bias.
reliability
Probabilistic metrics are computed at a grid of realistic lines, not at sportsbook prices. They measure whether the model's probabilities are calibrated. They do not measure profit.
| -7.0% |
| 455 |
| Hits allowed · MLB | edge rejected | 272–285 | +0.20pp | -8.7% | 49 |
| Hits · MLB | edge rejected | 2340–2664 | — | -3.4% | 0 |
| Strikeouts · MLB | earning | 278–274 | +0.21pp | -4.3% | 54 |
| Walks allowed · MLB | earning | 245–251 | +0.21pp | -6.5% | 33 |
| To win the fight · UFC | earning | 14–13 | +0.17pp | +27.4% | 29 |
| Total rounds · UFC | earning | 7–11 | +0.21pp | -26.1% | 19 |
| Total rounds · UFC | earning | 8–6 | +0.00pp | +5.5% | 14 |
| Receptions · NFL | earning | 1–3 | — | -49.5% | 0 |
| Rush attempts · NFL | earning | 15–15 | — | -8.1% | 0 |
| Rushing yards · NFL | earning | 41–41 | — | -5.9% | 0 |
| Interceptions · NFL | earning | 5–5 | — | -0.1% | 0 |
| Anytime TD · NFL | earning | 14–14 | — | +23.8% | 0 |
| Pass attempts · NFL | earning | 15–15 | — | -5.5% | 0 |
| Completions · NFL | earning | 13–13 | — | -4.8% | 0 |
| Passing TDs · NFL | earning | 28–28 | — | -1.1% | 0 |
| Passing yards · NFL | earning | 23–24 | — | -7.4% | 0 |
| Receiving yards · NFL | earning | 81–80 | — | -5.0% | 0 |
| lgbm |
| 0.634 |
| 0.689 |
| +7.9% |
| < 0.0005 |
| 0.5589 |
| 0.0224 |
| 0.932 |
| 3,230 |
extra points extra_points | lgbm | 0.759 | 0.817 | +7.1% | < 0.0005 | 0.5609 | 0.0334 | 0.921 | 3,110 |
Field goals made fg_made | lgbm | 0.677 | 0.724 | +6.5% | < 0.0005 | 0.5008 | 0.0220 | 0.934 | 3,110 |
First touchdown first_td | lgbm | 0.052 | 0.068 | +23.7% | < 0.0005 | 0.2088 | 0.0069 | 0.952 | 28,775 |
Sacks idp_sacks | lgbm | 0.221 | 0.237 | +6.9% | < 0.0005 | 0.3414 | 0.0094 | 0.948 | 11,561 |
Kicking points kicking_points | lgbm | 2.081 | 2.220 | +6.3% | < 0.0005 | 0.6091 | 0.0137 | 0.830 | 3,110 |
Last touchdown last_td | lgbm | 0.050 | 0.066 | +24.5% | < 0.0005 | 0.2035 | 0.0082 | 0.952 | 28,775 |
Pass attempts pass_att | lgbm | 5.545 | 6.433 | +13.8% | < 0.0005 | 0.5360 | 0.0243 | 0.852 | 3,417 |
Completions pass_cmp | lgbm | 3.819 | 4.404 | +13.3% | < 0.0005 | 0.5440 | 0.0219 | 0.831 | 3,417 |
Interceptions pass_int | lgbm | 0.421 | 0.466 | +9.6% | < 0.0005 | 0.4102 | 0.0228 | 0.926 | 3,417 |
pass longest pass_longest | lgbm | 7.323 | 9.139 | +19.9% | < 0.0005 | 0.5870 | 0.0187 | 0.759 | 3,297 |
Passing TDs pass_td | lgbm | 0.582 | 0.631 | +7.6% | < 0.0005 | 0.4356 | 0.0201 | 0.939 | 3,417 |
Passing yards pass_yds | lgbm | 43.891 | 52.196 | +15.9% | < 0.0005 | 0.5426 | 0.0173 | 0.808 | 3,417 |
points allowed points_allowed | lgbm | 5.305 | 5.953 | +10.9% | < 0.0005 | 0.6069 | 0.0136 | 0.843 | 3,230 |
Longest reception rec_longest | lgbm | 5.752 | 7.927 | +27.4% | < 0.0005 | 0.4857 | 0.0043 | 0.821 | 21,112 |
Rush + rec yards rec_rush_yds | lgbm | 16.448 | 18.250 | +9.9% | < 0.0005 | 0.4162 | 0.0098 | 0.826 | 26,800 |
Receiving yards rec_yds | lgbm | 13.907 | 15.493 | +10.2% | < 0.0005 | 0.4101 | 0.0102 | 0.844 | 24,424 |
Receptions receptions | lgbm | 1.033 | 1.137 | +9.1% | < 0.0005 | 0.3810 | 0.0055 | 0.906 | 24,424 |
Rush attempts rush_att | lgbm | 2.513 | 2.795 | +10.1% | < 0.0005 | 0.3227 | 0.0111 | 0.851 | 9,484 |
Longest rush rush_longest | lgbm | 4.849 | 6.164 | +21.3% | < 0.0005 | 0.5205 | 0.0165 | 0.744 | 9,048 |
Rushing yards rush_yds | lgbm | 14.937 | 16.243 | +8.0% | < 0.0005 | 0.4184 | 0.0081 | 0.812 | 9,540 |
solo tackles solo_tackles | lgbm | 0.896 | 0.956 | +6.3% | < 0.0005 | 0.4267 | 0.0042 | 0.910 | 36,678 |
Tackles + assists tackles_assists | lgbm | 1.243 | 1.314 | +5.4% | < 0.0005 | 0.3730 | 0.0033 | 0.888 | 36,678 |
pit
reliability
pit
reliability
pit
reliability
pit
reliability
pit
reliability
pit
reliability
pit
reliability
pit
reliability
pit
reliability
pit
reliability
pit
reliability
pit
reliability
pit
reliability
pit
reliability
pit
reliability
pit
reliability
pit
reliability
pit
reliability
pit
reliability
pit
reliability
pit
reliability
pit
reliability
pit
reliability
pit
reliability
pit