StatIQPregame to final — all in one place

Model Report Card

How our models score against real closing prices, sport by sport — including the sports where the closing price beats us. Every number cites the document it came from and the date it was measured.

How to read this page

  • BACKTEST means a model was run over past seasons and scored against archived prices. It is evidence about a method. It is not money, not a record, and never a promise.
  • LIVE means every play was published before the event started and graded afterwards, in the open, and never deleted.
  • Rates need at least 20 graded decisions before we print them. Under that they show a dash — in both directions.
  • Every number here cites the research document it came from and the date it was measured. A validator refuses to ship this page if a cited quote is no longer in its document.

Curated 2026-08-06 from the research documents listed at the bottom. Live records are read from the API at page load, never stored here.

Where we lose

The closing line beats our probability in NBA, NHL and MMA. Not narrowly, and not in one segment — everywhere we looked, in every season we scored. We publish this because a model that has never been measured against the closing price is not a model, it is a marketing asset.

  • NBABACKTEST

    On 193,342 real closing prices the no-vig close beat our probability in every market, in both seasons. We ran 119 segment tests hunting for one place we were ahead. Zero survived — noise alone should have handed us about three.

    NBA — the verdict layer at the REAL closing prices · measured 2026-08-06

  • NBA points + PRABACKTEST

    Once the false confidence is stripped out, our projection carries essentially no probability information on the two biggest NBA markets. The honest displayed number there is a near-constant 0.49.

    NBA — the verdict layer at the REAL closing prices · measured 2026-08-06

  • NHLBACKTEST

    On 164,099 real closing prices the market won every market out of sample. An internal note had us beating the market on saves at n=265; at scale it flipped sign, so we killed it. That correction is recorded in the research doc, not quietly deleted.

    NHL — close-line validation at scale · measured 2026-08-05

  • MMABACKTEST

    Against 2,680 real closing prices the market beat our fight-winner model in all six years, and its lead is widening — 2026 is our worst year. Flat-stake economics are negative at every threshold, at best price and at consensus, and the bigger our disagreement with the market, the worse we did.

    MMA (UFC) — fight-winner model vs the real closing market · measured 2026-08-06

  • CFBBACKTEST

    Five documented attempts to replace the shipping CFB model — v4, v5 (a quantile distribution layer), a calibration shrink, v6, and v7 (weather) — were all tested and all rejected. The research document calls the last one the sixth defense of a model we have never managed to beat.

    CFB — v3 evidence and the rejected successors · measured 2026-08-05

That is why we do not sell NBA, NHL or MMA picks; why our displayed probability is anchored to the market's own price instead of our model's confidence; and why the only things we claim in those sports are the price we found and the record we published.

Every sport, one row

SportScored againstSampleOur probability vs the closing priceLive record
MLBNo closing-price study committed for MLBNot measured437-313 · 58.3% · +11.2%per play, current model era since 2026-07-28
WNBANo closing-price study committed for WNBANot measured74-57-1 · 56.5% · +6.0%per play, current model era since 2026-07-28
College BasketballReal 2025-26 closes, one snapshot ~14 min pre-tip24,681 two-sided props scored for calibrationStatIQ ahead — Brier .24468 vs .24541None — not launched
NBAReal closes, 2024-25 + 2025-26 incl. playoffs193,342 prices scoredMarket ahead — pooled Brier gap +0.0117 (z +29.2)None — not launched
NHLReal closes, 2024-25 + 2025-26 incl. playoffs164,099 prices scoredMarket ahead — pooled Brier .2407 vs .2289 (z +32.5)None — not launched
MMA (UFC)Real closes, 2021-2026, six years walk-forward2,680 fights scoredMarket ahead — log-loss .642 vs .593None — not launched
College FootballClose-snapshot backtest prices — no closing-price archive scored476 backtest picks (2024-25 test window)Not measuredNone — backtest preview

4 of these 7 sports have been scored against a banked archive of real closing prices. 3 have not, and their rows say so rather than borrowing another sport’s verdict.

Where we are ahead — one sport, one number, still a backtest

College basketball is the only place in this study where our calibrated probability scores better than the market's own no-vig closing price. It is a backtest, it is one season, and it is the single result on this page we would call a finding.

  • CBBBACKTEST

    Our market-anchored hit probability beat the no-vig close on 24,681 two-sided props at the real 2025-26 closes: Brier .24468 vs .24541, log-loss .68223 vs .68372. Small margin, large sample, and it is the only sport where the sign points our way.

    CBB — real-price re-score at the banked 2025-26 closes · measured 2026-08-06

  • CBB · projection accuracyBACKTEST

    The v3 projector beat the shipping blend on MAE in all four markets on a season it had never seen: points +3.27%, rebounds +4.03%, assists +5.20%, threes +3.30%, on 58,266 test rows per market.

    CBB — real-price re-score at the banked 2025-26 closes · measured 2026-08-06

No live CBB record exists yet — the season starts in November. Until then this stays labelled BACKTEST, and we make no ROI claim from it.

Sport by sport

MLB

LIVE — MEASURINGLIVE RECORD

LIVE — measuring in public. No efficiency verdict published.

The only thing we claim about MLB is the record, and the record is on this page — computed from the live board at page load, not typed in.

The graded record — every play published before the event

LIVE
EraW-LHit rateReturn per playUnitsSample
Current modelsince 2026-07-28437-31358.3%+11.2%+83.77u750 decisions
Previous modelsince 2026-07-12354-32552.1%−2.0%−13.90u679 decisions
Whole official recordsince 2026-07-12791-63855.4%+4.9%+69.87u1429 decisions

A model cutover starts a new era; the previous era stays published and unedited, because a scoreboard you are allowed to reset is not a scoreboard. Return per play is flat 1-unit staking on the price we locked, before any bonus, boost or fee. It is what happened, not a forecast.

Source: the live MLB board · read at page load · graded through 2026-08-12 · same rows and same cutoffs as the Record tab

What we claimed vs what happened — current model era, 750 graded plays

LIVE

Under-claiming by 3.3 points — we said 55.0%, the record came in at 58.3% over 750 graded plays.

Probability we displayedPlaysMean claimActually hit
under 45%10641.4%46.2%
45–50%14947.5%51.0%
50–55%11952.2%58.8%
55–60%11257.4%57.1%
60% and up26464.8%67.4%
All75055.0%58.3%

Brier score 0.2380 on the probability we published (a coin flip scores 0.2500; lower is better). This probability is anchored to the market’s own no-vig price, so when it lands calibrated most of that calibration belongs to the market, not to our projection — the same anchoring the research documents below argue for.

Source: the live MLB board · computed at page load from the graded rows of the current model era

What this does not prove

  • There is no committed MLB close-archive study in this repository's research folder, so this card shows no backtest at all. What you see is the live graded record and the live calibration of the probability we displayed before each game.
  • The displayed probability is anchored to the market's own no-vig price. When it comes out calibrated, most of that calibration belongs to the market, not to our projection.
  • Rates are shown only once a bucket has at least 20 graded decisions. Below that they render as a dash, in both directions — a −100% at 0-4 and a +50% at 2-1 are equally fake.

What we shipTracker + best price + calibrated probability. Every play logged before first pitch, graded in the open, never deleted.

WNBA

LIVE — MEASURINGLIVE RECORD

LIVE — measuring in public, and currently over-claiming.

Our WNBA board has been claiming a higher hit probability than it has delivered. The gap is computed live below from the same rows the record is graded on.

Walk-forward projection study
25,917 train / 4,041 test rows per market, cache through 2026-07-22

The graded record — every play published before the event

LIVE
EraW-L-PHit rateReturn per playUnitsSample
Current modelsince 2026-07-2874-57-156.5%+6.0%+7.93u131 decisions
Previous modelsince 2026-07-0849-5845.8%−3.2%−3.47u107 decisions
Whole official recordsince 2026-07-08123-115-151.7%+1.9%+4.45u238 decisions

A model cutover starts a new era; the previous era stays published and unedited, because a scoreboard you are allowed to reset is not a scoreboard. Return per play is flat 1-unit staking on the price we locked, before any bonus, boost or fee. It is what happened, not a forecast.

Source: the live WNBA board · read at page load · graded through 2026-08-12 · same rows and same cutoffs as the Record tab

What we claimed vs what happened — current model era, 131 graded plays

LIVE

Level — we claimed 56.4% and delivered 56.5% over 131 graded plays.

Probability we displayedPlaysMean claimActually hit
under 45%137.6%under 20 decisions
45–50%2048.5%55.0%
50–55%4252.9%47.6%
55–60%3057.0%60.0%
60% and up3864.3%63.2%
All13156.4%56.5%

Brier score 0.2444 on the probability we published (a coin flip scores 0.2500; lower is better). This probability is anchored to the market’s own no-vig price, so when it lands calibrated most of that calibration belongs to the market, not to our projection — the same anchoring the research documents below argue for.

Source: the live WNBA board · computed at page load from the graded rows of the current model era

Two model-quality lanes we tested and closed — both negative

BACKTEST
ExperimentResultShipped?
LightGBM parameter sweep, 39 configsAll 39 land within ±0.1% pooled MAE of current; best is +0.03% (noise)No — lane closed
Opponent threes-defense featurethrees MAE +0.05%, points +0.04% — nothingNo — real but inert

Projection-quality only; no odds, no ROI. Both were audit findings we could have shipped as "new model improvements" and did not.

Source: WNBA — model research notes (audit batch) · measured 2026-08-05

What this does not prove

  • No WNBA close-price archive study has been run, so there is no Brier-vs-the-close number for WNBA on this page. NBA, NHL, CBB and MMA have one; WNBA does not, and we would rather show the hole than fill it with a proxy.
  • The record splits at a model cutover (2026-07-28). Both eras stay published; the current-era numbers are the ones our headlines use.
  • Seven straight information/architecture additions across sports have come back inert. That is the prior we now apply to any new feature idea, including our own.

What we shipTracker + best price + calibrated probability, with a projection-gap band and a placement-correlation warning.

Research document: WNBA — model research notes (audit batch) · measured 2026-08-05

College Basketball

MODEL AHEAD (BACKTEST)BACKTEST ONLY

BACKTEST — the one sport where our probability beat the closing price.

Walk-forward, scored at the real banked 2025-26 closing prices: our calibrated probability beats the no-vig close, and the selective tier's ROI confidence interval excludes zero. It is still a backtest.

Price archive
156 days · 1,288 prop events · 106,554 outcome rows · one tip−10min snapshot per event
Join quality
27,958 archived props → 88.5% matched to a projectable player; name-level resolution 99.7%
Books
betmgm, fanduel, betonlineag, williamhill_us, bovada, betrivers

No forward record. No live CBB record exists. The season starts in November 2026; everything above is a walk-forward backtest scored at archived closing prices.

Projection accuracy — v3 vs the shipping blend (MAE improvement, 58,266 test rows per market)

BACKTEST
Prop marketMAE improvement over the production blend
Points+3.27% better
Rebounds+4.03% better
Assists+5.20% better
Threes+3.30% better

On a season the model had never seen. The same four improvements appeared in the 2024-25 pilot, so the feature edge replicates out-of-sample.

Source: CBB — real-price re-score at the banked 2025-26 closes · measured 2026-08-06

Probability calibration vs the no-vig close — 24,681 two-sided props

BACKTEST
Metric (lower is better)Our anchored probabilityNo-vig closing priceAhead
Brier.24468.24541StatIQ
Log-loss.68223.68372StatIQ

The only sport in this study where that column reads StatIQ. The margin is small; the sample is not.

Source: CBB — real-price re-score at the banked 2025-26 closes · measured 2026-08-06

Selection tiers at the real closes — BACKTEST, not a record

BACKTEST
TierPicksHit rateROI at consensusROI at best priceUnits
SELECTIVE (v3)6264.5%+35.2%+36.4%+21.8u
VOLUME (blend)57848.1%+1.3%+2.0%+7.3u

Wilson 95% on the selective hit rate [52.1%, 75.3%]; bootstrap 95% on ROI [+9.6%, +59.6%]. Excludes zero — on n=62, at archived snapshot prices, fee-free, with limits unknown.

Source: CBB — real-price re-score at the banked 2025-26 closes · measured 2026-08-06

What we claimed vs what happened, on the same 62 backtest picks

BACKTEST
Mean claimed hit probabilityRealized hit rateRead
0.48264.5%We under-claimed — on 62 picks, which is not enough to call it skill

0.482 is the number we would actually publish next to a pick. We are not going to advertise 64.5% off 62 backtest plays.

Source: CBB — real-price re-score at the banked 2025-26 closes · measured 2026-08-06

What this does not prove

  • The Volume tier is statistically zero at real prices (+1.3% on n=578, 48.1% hit) and volume-threes is actively negative at −12.9% on n=159. It ships labelled "more plays / lower hit / higher variance" and it will never be marketed as making more money.
  • 56 of the 62 selective picks are unders at plus money, and 71% of them live on a single book's quote — thin markets where live execution means that book or no bet, and limits are unknown.
  • The 2024-25 pilot's +14.8% was partly a price-construction artifact: the pilot pooled prices from books quoting other lines. Corrected to at-the-line construction, the true pick rate is roughly 0.25% of quoted props — one to two picks on a full CBB slate day.
  • This is a BACKTEST on archived snapshot prices, fee-free, limits unknown. It is not a forward record and no public ROI claim comes from it.

What we shipNovember: two clearly-labelled selectivity tiers over the same board, each with its own separate append-only record, graded independently and never mixed.

Research document: CBB — real-price re-score at the banked 2025-26 closes · measured 2026-08-06

NBA

MARKET WINSBACKTEST ONLY

EFFICIENT — the market wins. We ship tracker + best price + projections, never +EV.

193,342 real closing prices, two seasons including playoffs. The market beat our probability in every market, in both seasons, in all 119 segment tests. This is the strongest confirmation the project has produced for any sport — against ourselves.

Closing prices scored
193,342 — points 40,420 · rebounds 39,675 · assists 37,440 · threes 35,752 · PRA 40,055
Price archive
489 ET day files · 2,648 events / 2,641 with props (99.7%) · median 8 books
Are these really closes?
Realized lag median 14.4 min before tip (p10 14.3, p90 14.4)
Name join
193,350 matched = 98.9% of two-sided rows = 100.0% of playable, zero unknown names

No forward record. No live NBA record exists. NBA has not launched; every number on this card is a walk-forward backtest scored against archived closing prices.

Projection accuracy — walk-forward MAE vs the realized stat (2024-25 quoted rows)

BACKTEST
Prop marketBacktest MAERecorded production evalRead
Points5.0925.122backtest matches production
Rebounds2.0582.050backtest matches production
Assists1.5151.555backtest matches production
PRA6.5506.509backtest matches production

This is a fidelity check, not a scoreboard: it proves the backtest reproduces the shipping projector. No MAE-versus-the-market's-own-line comparison has been run for NBA, so this page does not show one.

Source: NBA — the verdict layer at the REAL closing prices · measured 2026-08-06

Probability calibration vs the no-vig close — 2025-26 out-of-sample Brier

BACKTEST
Prop marketOur NB probabilityMarket better byzAhead
Points.26350.0140+15.0The close
Rebounds.26180.0143+15.8The close
Assists.25370.0081+9.8The close
Threes.24440.0033+4.9The close
PRA.26720.0176+16.9The close

Lower Brier is better. Pooled the gap is +0.0123 (z +31.4) in 2024-25 and +0.0117 (z +29.2) in 2025-26 — the market ahead in both. Per-market Platt recalibration closes most of it (pooled .2585 → .2480 against a no-vig .2468) without closing it.

Source: NBA — the verdict layer at the REAL closing prices · measured 2026-08-06

How much of our confidence survives recalibration (Platt slope b)

BACKTEST
Prop marketPlatt bWhat that means
Threes0.615keeps real information
Assists0.306keeps real information
Rebounds0.195keeps a little
Points0.030displays a near-constant ~0.49
PRA0.008displays a near-constant ~0.49

b is the fraction of our raw confidence that survives contact with reality. On points and PRA — the two biggest markets — the honest displayed probability is a coin flip. NHL's equivalents were 0.29–0.41 across the board.

Source: NBA — the verdict layer at the REAL closing prices · measured 2026-08-06

Reliability — our top-confidence decile, 2025-26

BACKTEST
Probability sourceClaimsActually hits
Raw negative-binomial (ours)0.7900.559
Legacy line×0.15 normal band0.7950.526
Poisson0.8710.534
The no-vig closing price0.6260.626

The market claims exactly what it delivers. Every raw model arm we tried claims far more than it delivers — which is why none of them is allowed to reach a user.

Source: NBA — the verdict layer at the REAL closing prices · measured 2026-08-06

The edge hunt — 119 segment tests

BACKTEST
CutResult
Market · model side · price bucket · line size · rest/B2B · home/away · role · sample · books quoting · hold · season phase · blowout proxy0 of 119 segments where we beat the market (z ≤ −1.96)
Best profitable bucket found (Platt, edge ≥ 0.10, 2025-26)+9.53% ± 2.13 on n=3,033 — and −2.87% ± 2.23 on the replication control. Opposite signs. Not an edge.

Chance alone would have produced roughly three nominal winners in 119 tests. We got none. This lane is closed.

Source: NBA — the verdict layer at the REAL closing prices · measured 2026-08-06

What repeated across both seasons has no model in it at all

BACKTEST
Rule2024-252025-26
Take the best quoted price when it beats the market's own consensus fair value by ≥ 2 points+4.89% ± 2.22 (n=2,517)+6.85% ± 1.84 (n=3,647)
Pure price-taking at a flat 0.50 probability−4.78% ± 0.54−1.97% ± 0.55

That is line shopping, not projection skill — a property of books disagreeing at the close, untested forward, and pre-registered as a hypothesis for 2026-27. It is not a StatIQ betting edge and we do not present it as one. It is the quantified case for best-price being the product.

Source: NBA — the verdict layer at the REAL closing prices · measured 2026-08-06

What this does not prove

  • The backtest substitutes a ~200-line numpy histogram GBDT for LightGBM (not installed on the research box) and retrains monthly instead of nightly — training data is staler, never fresher. The MAE fidelity table above is how we checked that the substitution changes nothing.
  • The game-log caches include preseason games, and production merges the same files, so the first fortnight of every season projects rotation minutes off preseason rotations. A known wart, disclosed, being fixed before October.
  • We may not make any +EV, edge or ROI claim for NBA — not for a market, not for a segment, not for threes, not for the +9.53% bucket.

What we shipOctober: tracker + best price + projections. Displayed probability = the no-vig close plus 0.10 × (our model − the close) — the only arm in the study that beat the plain close out-of-sample, and it beat it by a hair.

Research document: NBA — the verdict layer at the REAL closing prices · measured 2026-08-06

NHL

MARKET WINSBACKTEST ONLY

EFFICIENT — the market wins every market. Tracker + best price only, never +EV.

164,099 real closing prices across two seasons including playoffs. Our distribution work is real and measurable; the market's price is still the better probability everywhere, including the one market where a smaller sample said otherwise.

Closing prices scored
164,099 — sog 41.3K · points 45.8K · assists 45.3K · goals 26.4K · saves 5.3K
Price archive
537 ET days · 2,790 events · 6-8 US books · snapshots median 4.4 min inside the commence−10min request
Name join
98.8% of two-sided rows = 99.8% of playable; only 10 rows truly unmatched
Out-of-sample discipline
Every spec decision made on 2024-25; 2025-26 held untouched

No forward record. No live NHL record exists, and no NHL product has shipped. Everything on this card is a walk-forward backtest scored against archived closing prices.

Distribution choice — negative binomial vs Poisson (2025-26 out-of-sample Brier)

BACKTEST
Prop marketPoissonNegative binomialzBetter
Shots on goal.2565.2543−13.0NB
Points.2516.2507−8.4NB
Assists.2286.2272−13.3NB
Goals.2178.2160−11.1NB
Saves.2634.2554−7.6NB

A real, replicated modelling result at 21× the sample of the first check — count props want the negative binomial. It is also a result about us, not about the market.

Source: NHL — close-line validation at scale · measured 2026-08-05

Probability calibration vs the no-vig close — 2025-26 out-of-sample Brier

BACKTEST
Prop marketOur NB probabilityNo-vig closing priceAhead
Shots on goal.2543.2443The close
Points.2507.2379The close
Assists.2272.2147The close
Goals.2160.2029The close
Saves.2554.2498The close

Pooled: no-vig .2289 vs ours .2407, z +32.5. Per-market Platt recalibration pulls us to .2340 — better, still behind. A k=0.10 no-vig anchor adds nothing over the no-vig itself here.

Source: NHL — close-line validation at scale · measured 2026-08-05

The correction we published against ourselves

BACKTEST
WhenWhat we foundStatus
2026-08-05, n=265On saves our NB calibration slightly beat the market no-vig (.2459 vs .2493) — the only market where the model outscored the benchmarkRecorded, explicitly marked NOT actionable
Same day, at scale2024-25 gives a nominal, insignificant −.0018 (z −1.1, n=2,642) and 2025-26 flips to +.0057 (z +4.0, n=2,699)DEAD — selection noise

This is what a small-sample nugget looks like when you go and check it. We keep both entries in the research doc so the failure is as findable as the finding.

Source: NHL — close-line validation at scale · measured 2026-08-05

Reliability — what raw model confidence is worth

BACKTEST
Raw model saysReality
.16.27 to .37
.74.57 to .60

The raw logit keeps only about a third of its claimed sharpness (fitted Platt b ≈ .29–.41). The market prices lineup, power-play role and goalie news our game-log model cannot see, so residual dispersion alone under-states the uncertainty.

Source: NHL — close-line validation at scale · measured 2026-08-05

What this does not prove

  • No MAE or WAPE study has been published for NHL projections, so this card shows no projection-accuracy table. The NHL work to date is a probability study.
  • Goals ran hot in the low buckets out-of-sample after recalibration (predicted .18, actual .26; the base rate shifted between seasons). The spec says refit yearly, display rounded probabilities, and never lean on the goals tail.
  • A market with no banked closing rows does not ship a probability at all.

What we shipNot built yet. When it is: tracker + best price, with the model probability labelled as ours and never as +EV — negative binomial, one dispersion size per market, per-market Platt recalibration refit every season.

Research document: NHL — close-line validation at scale · measured 2026-08-05

MMA (UFC)

MARKET WINSBACKTEST ONLY

EFFICIENT — the close beat us in all six years. Tracker pattern only, never a pick product.

A from-scratch UFC winner model, walk-forward against 2,680 real closing prices from 2021 to 2026. It is honestly calibrated and genuinely better than a coin. It is also beaten by the market every single year, and the gap is growing.

Fight spine
8,808 UFC fights scraped (8,652 decided), 2,731 fighter profiles
Price archive
2,042 day files · 5,118 events with US moneyline at commence−15min · median 11 books per fight
Join
2,728 / 2,891 UFC fights 2021–26 = 94.4%; 48 draws/no-contests removed → 2,680 scored
Leakage canary
Shuffling the training labels collapses out-of-sample skill to a coin (log-loss 0.699, 46% accuracy)

No forward record. No live MMA record exists, and no MMA product has shipped. Everything on this card is a walk-forward backtest scored against archived closing prices.

Our model vs the no-vig close, year by year

BACKTEST
YearFightsLog-loss (ours)Log-loss (close)Accuracy (ours)Accuracy (close)Ahead
2021386.654.63462.4%64.3%The close
2022490.651.59660.2%68.0%The close
2023497.649.59862.8%68.2%The close
2024497.632.58566.2%70.6%The close
2025509.632.58964.8%68.0%The close
2026301.632.55162.8%73.4%The close
Pooled2,680.642.59363.3%68.6%The close

A coin scores 0.693, so the model is real. The gap's confidence interval is [+0.037, +0.059] — zero is nowhere near it — and 2026 is the widest year. The MMA close is getting sharper, not softer.

Source: MMA (UFC) — fight-winner model vs the real closing market · measured 2026-08-06

Is the model at least honest? Yes — it is calibrated, just coarse

BACKTEST
DecileModel saysObserved
Top0.75474.3%
Bottom0.25120.1%

Decile win rates track the model's probability monotonically. Being calibrated is not the same as being useful: the market is calibrated too, and sharper.

Source: MMA (UFC) — fight-winner model vs the real closing market · measured 2026-08-06

Flat-stake economics, pooled 2021–26 — negative at every threshold

BACKTEST
Bet when model − close ≥FightsROI at best priceROI excluding the exchangeROI at the median book
any edge2,680−3.8% [−9.2, +1.6]−4.3%−8.8% [−13.8, −3.8]
3 points2,241−3.3%−3.9%−8.6%
5 points1,942−4.7%−5.3%−10.1%
8 points1,562−5.9%−6.5%−11.4%
12 points1,052−6.7%−7.3%−12.7%

The more our model disagreed with the market, the worse it did — the classic signature of the market knowing something we do not: injuries, camp, weight cut, short-notice replacement. The "best price" column is flattered by an exchange we cannot use; the no-exchange column is the honest ceiling and it is still negative.

Source: MMA (UFC) — fight-winner model vs the real closing market · measured 2026-08-06

The hypothesis we went in with, refuted

BACKTEST
SegmentROI at 5-point threshold, best price
Prelims (bout 6 or later) — the "soft undercard" theory−8.4% [−16.6, −0.3]
Fights involving a debutant−17.7% [−32.6, −2.7]
Middleweight and up−21.4% [−32.4, −10.3]

We expected the market to win main events and lose undercards. The opposite is true: undercards are exactly where the market's private information is largest relative to what fight history can see.

Source: MMA (UFC) — fight-winner model vs the real closing market · measured 2026-08-06

What this does not prove

  • Around twenty segments were cut. A few came back positive with confidence intervals spanning zero — co-main +24.7%, five-rounders +16.3%, close fights +6.9%. Those are exactly the small-sample sparkles that have died every previous time we chased one, so they are recorded and not claimed.
  • A walk-forward blend of model and close beat the close in all five scoreable years, pooled gap −0.0020 with a confidence interval that crosses zero. The model carries a sliver the close has not priced. A sliver is not a product.
  • 2021's +6% at best price is a coverage-biased, main-card-skewed sample from the earliest market regime. 2026 is −17.5%. Do not chase the first number.

What we shipNot built. If MMA ever becomes a lane it is the NHL pattern: per-card winner probabilities with a band, best price on your books, an append-only record graded at the locked price, draws and no-contests voided. No EV framing anywhere.

Research document: MMA (UFC) — fight-winner model vs the real closing market · measured 2026-08-06

College Football

NOT MEASURED VS CLOSESBACKTEST ONLY

BACKTEST — never scored against a closing-price archive. Treat accordingly.

CFB has the oldest model program here and the weakest measurement. Its v3 projector beats the shipping blend on MAE and on backtest ROI — but no Brier-versus-the-close study has been run, so we cannot tell you whether the CFB market is efficient.

Walk-forward window
Train ≤ 2023, test 2024-25, identical policy on both arms — only the projection differs

No forward record. The CFB board runs as a labelled backtest preview until the season starts; the live status is read from the board itself below, so this page cannot claim a record the board does not have.

Board status right now: the CFB board is still a labelled backtest preview, so there is no live CFB record to show. Read from the board itself, not typed here.

v3 vs the production blend — BACKTEST at close-snapshot prices

BACKTEST
ArmPicksHit rateROIUnits
blend (previous production)28456.0%+11.9%+34u
v3 (launch projector)47657.1%+13.0%+62u

MAE: v3 beats the blend on every market — passing +1.4%, rushing +5.1%, receiving +5.5%. The v3-only picks ran +16.0% on n=287, so the extra volume carries information rather than noise.

Source: CFB — v3 evidence and the rejected successors · measured 2026-08-05

Everything we tried to replace v3 with, and what happened

BACKTEST
AttemptVerdict
v4 — new information lanesDid not clear the bar
v5 — quantile distribution layerDid not clear the bar
Calibration shrinkOverconfidence real but inert — no fix ships
v6 — three information lanesDid not clear the bar
v7 — weatherMechanism real, feature inert

Five rebuilds, five rejections, one shipping model. Every one of them would have made a better press release than a product.

Source: CFB — v3 evidence and the rejected successors · measured 2026-08-05

What this does not prove

  • Close-snapshot, fee-free, sane-best prices: the comparison between arms is clean, the absolute ROI level is optimistic. This is the research doc's own wording, not a disclaimer we bolted on.
  • No closing-price archive has been scored for CFB, so there is no Brier or log-loss number against the no-vig close on this page. Until that runs, treat CFB's market like NBA's — assume it is efficient — rather than assuming the backtest ROI transfers.
  • The passing-market wobble between arms is about four wins at n≈86. Noise.

What we shipv3 + the NB band, with the board's price construction restricted to books quoting at the consensus line.

Research document: CFB — v3 evidence and the rejected successors · measured 2026-08-05

Where these numbers live

  • NBA — the verdict layer at the REAL closing prices · measured 2026-08-06
  • NHL — close-line validation at scale · measured 2026-08-05
  • CBB — real-price re-score at the banked 2025-26 closes · measured 2026-08-06
  • MMA (UFC) — fight-winner model vs the real closing market · measured 2026-08-06
  • CFB — v3 evidence and the rejected successors · measured 2026-08-05
  • WNBA — model research notes (audit batch) · measured 2026-08-05
  • WNBAthe live WNBA board, read live on every page load, never typed by hand
  • MLBthe live MLB board, read live on every page load, never typed by hand

Every curated number on this page is validated before it can ship: the build fails if a claim quotes a sentence that is no longer in the document it cites, or cites a study date that document does not carry. The same picks, graded, are on the Record; the method and our own errata are on How we measure.