Every number on this site, where it comes from, how it was checked, and where it stops being trustworthy. The figures below are measured, not illustrative. If you read one section, read the last one.
Three independent sources, each checked against the others rather than trusted.
Play-by-play and rosters come from nflverse, the open NFL data project: 484,254 plays and ten seasons, 2016 through 2025. That becomes 57,737 player-weeks with 350 features each, every one a game that was played. A further 15,487 rows describe fixtures that have NOT been played — one per rostered player per remaining game — so the model can project the match that is next rather than the last one it has statistics for. Those rows carry no outcome, and the training code refuses any row without one.
Prices come from SportsGameOdds for player props and ESPN for game lines. Both are polled and archived, so the opening price is kept alongside the current one — that is what makes closing line value measurable at all.
Rosters are refreshed on a schedule that follows the football calendar: daily from August through November, when cutdowns take squads from 90 to 53 and the trade deadline lands, weekly the rest of the year. All 2,905 players on 2026 rosters are verified against the source after every sync — zero on the wrong team, zero missing, zero no longer rostered. Retired and released players are excluded rather than listed: the roster answers "who can appear in a market", and fourteen who could not were removed, Adam Thielen among them.
That check exists because it once failed badly. Before it, every player was filed under the team he finished the previous season with: 461 had changed team, 896 were no longer rostered anywhere, and 682 rookies and signings were missing entirely.
A price that reaches this site has passed four checks. Most do not.
implied(over) + implied(under) > 1.0 + ε
A sportsbook always holds. Two sides whose implied probabilities sum to exactly 1.0 describe a market with zero vig, which no book offers — that is a placeholder, not a price. Both sides quoted at +100 is the shape that produces it.
How it is verified: Of 384 two-sided props sampled from the live feed, 185 had an overround of exactly 1.0000 and only 37 held anything at all. Priced against those placeholders the board produced 27% edges, because de-vigging a zero-vig market returns 0.5 and the “edge” becomes the model's own probability minus a half. They are now filtered out before anything is computed.
−100 < american < +100 → rejected
No American price can sit between −100 and +100. A feed returning them is not returning odds.
How it is verified: One provider's free tier returned 1,593 props for a single week where 100% of payouts fell in that impossible range, flagged in its own payload as Scrambled, with every quarterback's passing line within fifteen of 100. That source is not used.
commence_time > now
“Current odds” means odds on a game you can still bet. A row for a finished game is not stale data, it is wrong data, and it renders identically to a live one.
Seventeen models. Each one predicts a distribution, not a number, and none is scored on data it has seen.
Gradient-boosted quantile regression. Each market fits a set of models across the distribution — the 5th percentile, the 25th, the median, and so on — so the output is a full shape rather than a point. LightGBM, chosen by measurement: six learners were run over identical folds and it finished 0.26 percentage points behind the best for a third of the training time, and stacking them was worse than the best single model while costing twice as much again.
Why a distribution matters. At a line of 70, a tight distribution centred on 67 is a clear under and a flat one averaging 67 is close to a coin flip. Both print 67. That is why a player page shows a histogram: the bet is on the shape, and the average hides it.
Counts are modelled as counts. Receptions and touchdowns use a hurdle model with the zero mass fitted explicitly, because 22% of eligible receiving weeks land on exactly zero and a continuous model has no way to represent that.
Calibration. A raw model probability is not a betting probability. Four calibrators compete on held-out data — none, Platt, isotonic and beta — and the winner is recorded on the model card. Six of the seventeen markets use beta or Platt; the rest measured no improvement and use the raw output, which is the honest outcome when calibration has nothing to fix.
Validation is walk-forward. Fit on every season before season S, predict season S, move on. Nothing is ever scored on data it trained on. Significance comes from a paired block bootstrap over weeks, 2,000 resamples, paired because both models score the same rows and blocked because a week is the independent unit.
The intervals cover what they claim to cover. This is the check worth caring about, and it is easy to skip because a model can score well on the centre of its forecast while being badly wrong about its own uncertainty. Measured on rushing yards across 25 walk-forward folds and 6,988 rows the model never trained on:
| Stated interval | Actually contained the outcome | Error |
|---|---|---|
Each number on a card, what it is computed from, and what it means when it looks extreme.
P(outcome > line) from the fitted distribution, then calibrated
The model's own answer for this player, this week, against this specific number. It reads the full forecast distribution at the line, so a line sitting in a fat part of the distribution and one sitting in a thin tail produce very different probabilities from the same projection.
How it is verified: Walk-forward only. The probability you see comes from a model whose accuracy was measured on seasons it never trained on.
devig(implied(over), implied(under)) · Shin's method
What the book is really offering once its margin is removed. A −110 / −110 market implies 52.4% on each side, which sums to 104.8%; the 4.8% is the hold. Removing it proportionally is naive, so Shin's method is used, which accounts for the book shading against informed money. Every edge on this site is measured against this number, never the raw price, because the raw price always favours the book.
model probability − fair probability
How much more often the model thinks this hits than the de-vigged price implies. Bigger is not better past a point. Above 10% on a liquid market, a disagreement is far more likely to be a broken model than a mispriced line, and nothing above that ceiling is ever recommended or used in a parlay — the site refuses to trust its own outliers.
A real case from this board, and why the honest answer was neither the model's nor the book's.
The board showed Sam Darnold, rushing yards, over 2.5, at −117. The model gave it 77%. The de-vigged price gave it 50%. That is a 27% edge, which would be the largest on the board by some distance.
A large edge is exactly what both a mispriced market and a broken model look like, so the edge itself could not say which this was. Three things were checked.
Was the market mapped correctly? Yes — the feed literally says “Sam Darnold Rushing Yards Over/Under” at 2.5, with a genuine −117 hold on both sides. Not a placeholder, not a mislabelled market.
Was the model personalising, or projecting an average? Personalising. It separated Darnold at 12.6 rushing yards from Jaxson Dart at 38.1, and passing touchdowns from 0.91 to 1.89 across seven quarterbacks.
What has he actually done? 24 of his last 40 games cleared 2.5. Sixty percent. His last twenty totals were 5, 9, 0, 9, 2, 7, 5, 23, 0, −1, 11, −2, 0, 1, 2, 0, 24, 0, 0, 14.
So the model at 77% overstates, the book at 50% understates, and the truth sits between them. There is value on that over — worth roughly ten points, not twenty-seven. Neither side was simply right, which is why a check that could only vote for one of them would have been useless.
The pick is shown, flagged high risk for an implausible edge, and not recommended. That is the system working: it found something real, refused to overstate it, and showed you the evidence to judge for yourself.
Most parlays are bad bets. The one legitimate edge is correlation, and it is smaller than people think.
The vig compounds. Four legs at −110 hold roughly 17% against you, so every leg must be clearly positive before the combination breaks even. The builder states this on any parlay of four or more.
Where the edge is. Books mostly price parlays by multiplying the legs, as if a quarterback's passing yards and his own receiver's receiving yards were unrelated. They are not. This site measured 368 correlations from 57,737 player-weeks using Spearman rank correlation — rank rather than Pearson, because the marginals are skewed and zero-inflated and Pearson would be dominated by the tails.
The quarterback-to-receiver stack everyone reaches for measures +0.168, on 53,279 observations. Real, and much smaller than its reputation. Every result shows true chance beside if independent so the gap is visible rather than hidden inside one adjusted number.
How the joint probability is computed. A Gaussian copula: each leg's probability is mapped to a normal quantile, the measured correlation is imposed, and the joint is integrated by Monte Carlo. Exact answers in more than two dimensions have no closed form. The implementation was cross-checked against an independent one — every pairwise correlation matching to 1e-12 and joint probabilities agreeing within three standard errors — and against exact bivariate orthant probabilities where those exist.
Correlation is subtracted in the ranking even when it is positive. That looks wrong and is not. Positive correlation genuinely raises the joint probability, and the expected-value term already collects that upside. It also concentrates risk: the legs win together and lose together, so the bet is closer to one large wager than to several small ones. The penalty prices the concentration.
Four refusals, never overridden: negative expected value, a player's over combined with his own under, the same player twice, and any leg above the plausibility ceiling.
Below six legs every combination is examined. Above that the space is too large, so a guided search takes over and the result says so — “best of the combinations examined” and “best combination” are different claims.
It leads with CLV rather than profit, and that ordering is the whole argument.
Over a season of prop bets, variance is large relative to any realistic edge. You can be up money with a broken model and down money with a sound one, and twenty settled bets cannot distinguish those. A win rate that early is close to meaningless.
Closing line value asks a better question: was the price you took better than the one the market settled on? A bettor who consistently beats the close is being paid for information. One who loses to it is paying for the privilege, whatever this month's results say.
It is expressed in points of implied probability rather than in cents, because ten cents means something different at −110 than at +900.
The summary refuses to characterise a record too short to characterise. It earns the right to a verdict at 15 bets on CLV, where it needs 30 on results, and below that it says so rather than printing a confident percentage.
Closing prices are typed in by hand, because no free source publishes a feed of settled closing odds and generating them would make the only trustworthy number on the page fictional. Everything is stored in your browser, not in an account.
Ranked recommendations, each with reasons, risks, a confidence breakdown showing which components had data, and the player's own record. Rejected picks appear below the recommended ones with the reason attached — a hidden row teaches nothing, and reading why something was refused is the fastest way to learn the shape of a bad bet.
State a shape — 2 to 10 legs, four risk styles, same-game or cross-game — and it searches combinations and returns the best few. Below it, a manual slip prices anything you assemble yourself, with legs above the plausibility ceiling marked so the two halves of the page cannot disagree silently.
Everything currently priced in one sortable table: model projection, line, edge, expected value, stake. Rows failing a review gate are shown rather than hidden, with their flags.
All 2,919 players on 2026 rosters, every position. A player page carries game logs, season stats, usage by situation, splits by venue and roof and kickoff slot, injury history, contract, advanced metrics, and the model's full forecast distribution with a draggable line.
Season-long totals with the reasoning stated: base rate, age adjustment, workload trend, and the games-played estimate. That last one is now derived from games actually played rather than from the injury report, which never records season-ending injuries — before the fix, availability correlated −0.140 with reality and 73 players who played zero games were projected for a full 17. It now correlates +0.867.
The most important section. Nothing below is estimated or approximated — it is declared missing.
Sharp money % and public betting %. These come from paid services that instrument sportsbook ticket counts. No free feed carries them and nothing in play-by-play implies them. They are named as unavailable on every card rather than approximated, because a fabricated “78% sharp” is indistinguishable from a real one and would move your stake.
PFF grades, DVOA, block win rate. Proprietary and paywalled. Each confidence component lists which of its specified inputs are missing, so a thin score is visible as thin rather than silently weighted to zero.
Charted routes run. The participation feed carries one route per play — the targeted receiver's — so counting it per player would silently measure targets instead. The proxy used here counts presence on a dropback and is named for that rather than for what it approximates.
Game lines. No spread, total or moneyline carries a recommendation. A model was built and walk-forwarded against the closing spread and it lost, by 1.67%, at p = 0.9995. Over 2,639 games the closing line's average error on the home margin is 0.04 points. The market is that good. No badge is shown because none is earned.
And variance, which is not a software limitation. A realistic edge on a player prop is a few percentage points. Over a season that is a meaningful advantage; over ten bets it is invisible. Everything here can be right about the long run and badly wrong about your Sunday.
The system is built to be able to say no. On a typical slate it recommends nothing, and that is the correct answer for a liquid market rather than a fault. A tool that always finds a pick is a tool that has stopped measuring.
None of this is financial advice and none of it is a promise about any individual bet. It is a set of distributions with stated uncertainty, priced against real markets, showing its working so that you can disagree with it.
How it is verified: Enforced on every sync. It once caught 2,723 rows of the previous season's completed fixtures being served as upcoming, with full prices, because a scheduling endpoint answers for the last finished season when asked for a week without naming a year.
In-play pricing is a different question from pregame and must never sit beside it.
How it is verified: A feed presenting itself as a second sportsbook was quoting Detroit at home +1500 against the pregame −120. Across 168 moneylines quoted by both, the median gap was 1,640 cents. It is excluded.
| 50% |
| 50.3% |
| +0.3 |
| 80% | 80.1% | +0.1 |
| 90% | 90.4% | +0.4 |
Every level lands within half a percentage point, slightly wide rather than slightly narrow, which is the safe side to err on. A split-conformal correction is fitted on held-out data as a backstop and applied where it is needed; on these markets it comes out at zero, because there is nothing to correct.
Every market, its CRPS, and how far it beat a naive recent-form baseline — read from the measurement each promoted model recorded at promotion, so a retrain moves this table rather than dating it. Markets typically report p < 0.0005, and that is the FLOOR of the test rather than a measurement: the estimator is (b+1)/(B+1), so with 2,000 resamples the smallest attainable value is 1/2001. Two markets showing the same figure cannot be compared on it; the CRPS gain beside it is the number that separates them.
| Market | CRPS | vs baseline | Training rows |
|---|---|---|---|
| Longest reception | 5.752 | +27.4% | 21,112 |
| Last touchdown | 0.050 | +24.5% | 28,775 |
| First touchdown | 0.052 | +23.7% | 28,775 |
| Longest rush | 4.849 | +21.3% | 9,048 |
| Pass longest | 7.323 | +19.9% | 3,297 |
| Passing yards | 43.89 | +15.9% | 3,417 |
| Pass attempts | 5.545 | +13.8% | 3,417 |
| Completions | 3.819 | +13.3% | 3,417 |
| Points allowed | 5.305 | +10.9% | 3,230 |
| Receiving yards | 13.91 | +10.2% | 24,424 |
| Rush attempts | 2.513 | +10.1% | 9,484 |
| Rush + rec yards | 16.45 | +9.9% | 26,800 |
| Interceptions | 0.421 | +9.6% | 3,417 |
| Receptions | 1.033 | +9.1% | 24,424 |
| Anytime TD | 0.167 | +9.0% | 28,938 |
| Rushing yards | 14.94 | +8.0% | 9,540 |
| Dst takeaways | 0.634 | +7.9% | 3,230 |
| Passing TDs | 0.582 | +7.6% | 3,417 |
| Extra points | 0.759 | +7.1% | 3,110 |
| Sacks | 0.221 | +6.9% | 11,561 |
| Field goals made | 0.677 | +6.5% | 3,110 |
| Kicking points | 2.081 | +6.3% | 3,110 |
| Solo tackles | 0.896 | +6.3% | 36,678 |
| Dst sacks | 0.958 | +6.0% | 3,230 |
| Tackles + assists | 1.243 | +5.4% | 36,678 |
CRPS is the continuous ranked probability score: it rewards a distribution for being both accurate and appropriately confident. Lower is better, and it is in the units of the stat, so a figure near fourteen on receiving yards means roughly fourteen yards of average distributional error. The percentage is what matters across markets, since the raw numbers are not comparable.
How it is verified: The ceiling earned itself. On the first board built from real prices, 16 of 26 legs sat above it, and the one investigated in detail turned out to be model error, not value.
Σ(component × weight) ÷ Σ(weight of components that had data)
Ten weighted components: team strength, quarterback, offensive and defensive matchup, injury, coaching, situational, market, player matchup, historical. The small percentage under the ring is coverage — how much of that weight had real data behind it. A 70 at 95% coverage and a 70 at 40% are entirely different claims, so the ring is coloured by coverage rather than by score and they cannot be mistaken for each other. Below 60% coverage nothing is recommended, however good the score looks. A row with no data returns insufficient data rather than a low number.
The weights are measured, not asserted. They started as a specification — quarterback 0.15, coaching 0.05, numbers someone wrote down. They are now fitted on 99,445 settled player-games, asking which components actually predict that a projection landed close. Fitted through 2023, with 2024 and 2025 held out entirely: 57.5% accuracy against a 51.3% base rate on 20,605 rows the fit never saw. That moved player matchup from 0.10 to 0.45 and situational from 0.10 to 0.007 — who a player faces carries most of the information, and rest days and travel carry almost none. What this does not show is profit: it fits projection accuracy, not return at a price, because no bet here has settled yet.
count(games where the player cleared this line) ÷ games
A plain tally from the player's own game log, up to 40 games. It shares no assumptions with the model, which is exactly why it is worth having: a second model would agree with the first for the same wrong reasons. A record that contradicts the model on a sample of 25 or more blocks the recommendation outright, whatever the edge says. Below five games it declines to have an opinion.
p × (decimal − 1) − (1 − p)
Profit per unit staked if the model probability is right. +6.5% means that for every $100 staked, this bet returns $6.50 on average over many repetitions — not that you will win 6.5% more often, and not a promise about any single bet.
min( kelly × 0.25 , 0.02 ) × bankroll
Quarter Kelly, capped at 2% of bankroll. Full Kelly maximises long-run growth if the probability is exactly right; a model probability is an estimate with real error, and full Kelly on a slightly wrong number is ruinous. The quarter is the conventional discount for that uncertainty and the hard cap is what stops a single confident wrong number from mattering.
It sizes ONE bet, and the parlay page shows five alternatives. Those five are near-misses of each other and usually share most of their legs, so placing all of them is not five independent bets at their stated stakes — it is several times the intended exposure to the same handful of players. The page warns when they overlap and shows what the combined stake would actually be, and each card reads “stake if taken alone” for that reason.
The bankroll box starts empty on purpose. The model computes a fraction, never a dollar amount, and until you enter a bankroll that is exactly what is shown — “0.29% of bankroll” rather than a figure in pounds or dollars. Earlier versions pre-filled $1,000 in three separate places, which was not a recommended bankroll but read as one; someone with $200 was doing arithmetic against a number this site invented. Enter your own and it is remembered in this browser only, shared across the pick cards, the parlay page and the manual builder, and never sent anywhere.
How it is verified: Checked against the closed form. Kelly is (pd − 1) ÷ (d − 1), which for a parlay at decimal 6.38 and joint probability 16.662% gives 0.011720; the implementation returns 0.011720, so the fraction staked is min(0.011720 × 0.25, 0.02) = 0.29% of bankroll — $2.93 against a $1,000 bankroll, and correspondingly less against a smaller one. Expected value matches p(d−1) − (1−p) to six places.
current price − opening price
A price that has shortened since it opened is money arriving on that side. The projection is weak and labelled as such: real closing line value needs the closing price, which by definition does not exist before kickoff.
Computed from dispersion and doubt, never from edge size: how wide the forecast is relative to its mean, whether the player carries an injury designation, how thin his history is, how much of the confidence model had data, and whether his record contradicts the model. A pick can carry a large edge and a high risk rating at once, and usually should — a big edge on a wide distribution is what model error looks like.
edge ≥ 2% AND edge < 10% AND EV > 0 AND coverage ≥ 60% AND no flags AND record not contradicting
Every clause has to hold. Failing any one is enough to drop a pick out of the recommended list, and each clause was added because something got through without it.
How every model scored on data it had never seen, and exactly which version is live — including the git commit it was trained from and whether the working tree was clean at the time.
Both kept in your browser rather than on your account by default — a bet log and a star list are private, and neither needs a round trip. The tracker can be synced to your account when you want it counted in your personal record; the watchlist stays local, so it does not follow you to another device.