StatIQNFL projections — with the receipts

How we measure

Every claim we make, the query behind it, and the result that would prove it wrong — including the mistakes we have made and corrected.

This is the same document we work from internally, published unedited. The numbers in it are regenerated from the live boards, not typed by hand, so it cannot quietly fall out of date. If you want the underlying rows, the audit log itemises every locked price and its close, and exports to CSV so you can recompute the totals yourself. The record is the same picks, graded.

1 section withheld from this page. We would rather tell you that than quietly publish a shorter document. Reason: this claim's own position is that we say nothing publicly yet, so publishing it would contradict it; the finding is also capacity-limited and would be destroyed by an audience.

StatIQ Claims Ledger

Every public claim we make, the query that backs it, and the result that would prove it wrong.

This exists so that when someone asks "why do you say that," the answer is a number and a method, not an opinion. Others are free to disagree — but they have to disagree with the data, and they can check it themselves.

Three rules for this document:

  1. The falsifier is written before the numbers come in, not after. "That metric doesn't apply to us" is a legitimate methodological position and also the exact sentence a losing bettor reaches for. The only thing that separates them is whether the test was specified in advance. If we ever want to drop a claim, we change it here first, in the open.
  2. The unflattering entries stay. A ledger you would not hand to a skeptic is not a ledger. The entries below where we measured against ourselves are what make the rest credible.
  3. Numbers are generated, never typed. Every figure in the tables below is read straight off the live boards by a script and written into this document mechanically. Nobody types a number into it, so it cannot quietly fall out of date and it cannot be nudged.

Claim 1 — Our all-time MLB record has been profitable. We report it as history, not a forecast.

What we say: the all-time official MLB tracked record — every pick logged before the game and graded at the price we locked, both model eras included — has been profitable: +308.41u over 3,722 graded plays, +8.29%, since 2026-07-12. It is a measured past result at a flat 1 unit per play, before any bonus, boost or fee, with pushes counted as no gain and voids excluded — a tracked return at recorded prices, not executed wagers, so not proof that any given bettor could have taken those prices. It reflects the whole selection, price-shopping and locking process, not model skill alone. It is not a per-bet +EV claim and not a promise about future picks.

Scope, stated plainly. MLB, all-time, both model eras — the current model (since 2026-07-28) and the previous model before it — never blended with another sport. WNBA's tracked record is +5.62% over 696 graded plays, positive but a smaller sample, reported the same way. NFL 2026 is a measurement-only forward test at −7.7% — a loss, published at equal weight, on which no money was ever wagered.

Why a big sample matters, and how we test significance. A single play's return has a standard deviation near one unit, so a few hundred plays cannot separate a real result from variance. The significance column below reports each sport's distance from zero using a standard error clustered by game-day — picks on the same slate share exposure (one game runs hot or cold; a total moves a whole slate the same way), and counting them as independent overstates significance. Clustered, MLB is +4.2σ (a naive independent-play calculation would read +5.1σ); WNBA is +1.4σ.

The bar we pre-registered (2026-07-20). We committed, in writing before the numbers matured, to make a forward profitability claim only once ROI clears 2σ from zero on n ≥ 1,500 tracked plays. MLB's clustered significance is past that bar; WNBA's is not. We publish the MLB history above either way — a graded record is a fact — and hold any forward-looking claim to that pre-registered standard.

Status: ACTIVE.


Claim 2 — CLV is a diagnostic in WNBA/MLB props, not a pass/fail gate

What we say: we publish our closing-line value, and we do not treat negative CLV in these markets as evidence we are wrong.

Why: CLV is a proxy. Its entire validity rests on the premise that the closing line approximates the true price — which requires informed money to move it there. We tested that premise on our own board and it failed:

  • Stratified close-vs-open AUC gap ≈ +0.008 (see table). The close predicts outcomes no better than the open. Stratification by market × line × side is mandatory; pooled, this test measures line level and reports a large fake gap.
  • Matched line-move test: comparing moved vs un-moved lines within the same market and same open line, movement produced no lift (up-moves −0.5 pts, −0.1σ; down-moves −3.9 pts, −0.6σ).
  • What looked like movement mostly wasn't. 184 of 251 MLB "moves" were exactly 1.0 on markets whose only lines are 0.5 and 1.5 — alt-line relisting, not a market moving.

A proxy that fails its own premise does not overrule a direct measurement. ROI is the primary measure. CLV is published as a diagnostic with this caveat attached.

Scope limit — this does not travel. It applies to thin prop markets we have measured. For NBA and NFL sides the close is very likely informative, and CLV remains a gate there.

Falsifier: if the stratified AUC gap in these markets reaches +0.03 or more at n > 5,000, the close is carrying information, and CLV returns to being a gate.

PINNACLE: MEASURED 2026-07-20, and the argument we were making is REFUTED. We had claimed "Pinnacle prices these props at ~7% vs ~2% on core markets, so it declines to stand behind its number and no sharp reference exists." We had never pulled it — our production request uses regions us,us2,us_ex,us_dfs and Pinnacle lives in eu. We went and pulled it directly, and that settles it:

PinnacleUS field (same events, same moment)
WNBA props6.96% (28 legs)6.41% – 8.85% across 9 books
MLB props6.99% (40 legs)6.38% (BetOnline)

The 7% number was right. The inference from it was wrong. Pinnacle sits squarely in the middle of the field — tighter than DraftKings (7.53%), ESPNBet (8.78%) and BetParx (8.85%), looser than BetOnline (6.41%) and Novig (6.44%). ~7% is simply what player props cost across the entire industry; it is not Pinnacle declining to price them. It prices them, on both sports, at market-typical margin.

Which means the hold argument is a dead end in BOTH directions. Margin tells us about a book's pricing model, not about the accuracy of its number. We cannot conclude "no sharp reference exists" from vig, and we equally cannot conclude Pinnacle is sharp from vig.

The real test is still undone: score our graded picks against Pinnacle's de-vigged number and ask whether it predicts outcomes better than our consensus close does. That is the only thing that would confirm or kill Claim 2, and it requires either historical Pinnacle odds (billed 10×) or forward collection from now. Until it runs, Claim 2 rests on the stratified AUC test and the matched-move test only — and the Pinnacle line must not be cited as supporting evidence.

Status: ACTIVE (established 2026-07-20).


Claim 3 — Line-shopping and lock discipline cut the hold we pay by ~64%

What we say: taking the best available price across books materially lowers the edge required to break even, and this is true whether or not our projections are any good.

Why: it is mechanical, not predictive. This is the most defensible thing we do — it requires no model to be right, and it is verifiable on any single pick. Measured per pick from the per-book ladders at a single instant (MLB, n=1,011):

median single book   7.29%
best single book     5.95%
line-shopped         4.29%     <- shopping buys 3.00 pts vs median, 1.67 pts vs best book
after our lock rule  2.65%     <- waiting for a mature market buys another 1.64 pts

Two distinct mechanisms — shopping and lock discipline — each separately measurable. Combined, 7.29% → 2.65% is a 64% reduction in the hold, i.e. break-even edge falls from 3.6 pts to 1.3 pts.

Honest caveats, stated up front:

  1. Quote the vs-BEST-BOOK number (1.67 pts), not the vs-median one (3.00). A user who simply bets at the sharpest book they already have captures most of the benefit with no tool at all. The median-book comparison flatters us by assuming the user bets randomly.
  2. The shopped figure is a synthetic cross-book pair — what shopping buys, not a two-way market anyone could bet both sides of.
  3. The "after lock rule" figure is conditioned on the lock rule itself (locking requires hold below a threshold). It is what we pay; it is not an achievement of shopping.
  4. We have no WNBA numbers here. WNBA stores per-book prices for our side only, so a genuine single-book hold cannot be computed from stored data. This claim is currently MLB-only.

Correction on record (2026-07-20): the first version of this table reported "3.56% → 1.75%" and was wrong twice, both times in our favour. It used close_over/close_under as "a book's hold" when those are already best-of-N across books, and it measured the shopped column on locked picks only — selecting on the very threshold that defines locking. Real single-book holds are 7–9%, not 3.5%.

Falsifier: if line-shopped hold stops beating the BEST SINGLE BOOK by at least 1 point — because we lose book access, or because the books we can reach converge — this claim dies with it. Tracked in the generated table below, which reports both comparisons every run.

Status: ACTIVE.


Claim 4 — The prop markets we have measured are efficient, and we say so

What we say: WNBA and MLB player props measured efficient against our model. We do not claim to have found a persistent edge in them.

Why: this is the entry we would most like to delete, which is exactly why it stays. Repeatedly through 2026, apparent edges found at n < 150 dissolved at larger n: the blowout-under thread (full season, 302 games, −5.31%), "unders survive" (noise, +1.2σ, calibration sign reversed), the WNBA totals regime story (+1.9σ in one season, −1.3σ in the other, +0.1σ combined). The model is frozen against tuning until at least 2026-07-25 for this reason — we are collecting data, not fitting it.

Falsifier / graduation: an edge is real when it clears 2σ on a pre-registered, forward-tested sample — never on the sample that suggested it.

Status: ACTIVE.


Claim 6 — We publish what PROVES us, not everything that BUILDS us

What we say: our record, our grading, our method and our mistakes are open and checkable. Our per-pick model inputs are not a permanent public commitment.

Why this needs deciding in advance. Everything we publish is public JSON: 8,323 MLB and 1,953 WNBA picks, each carrying proj, hit_prob, ev, edge, the line, and the graded actual. That is a labelled dataset of our model's outputs. Anyone competent with an AI assistant can scrape it and train a surrogate that reproduces our projections to within noise — no code, no injury feed, no minutes model required. A subscriber could pay for one month and walk out with a working clone. This is not a hypothetical; it is a weekend of work for a technical person.

The distinction that resolves it:

  • Numbers that PROVE us — every pick logged before tip, the result, the price we took, the closing price, CLV, ROI with error bars, the void/DNP reasons, this ledger and its errata. These stay open. They are the entire basis of the claim that we are honest, and hiding any of them would forfeit the one position nobody else in the category occupies.
  • Numbers that BUILD us — the per-pick projection and hit probability on pending picks, i.e. the live model output. These are open today by default, not by decision, and that default is what a cloner needs most.

Current exposure is genuinely zero, on two independent grounds: no one has signed up yet, and our own measurements say the projection adds ~nothing beyond the market ([[Claim 4]]) — so a clone of it reproduces a model that measured efficient. Both of those change the moment either one changes.

TWO DIFFERENT THREATS. The first version of this claim conflated them and proposed a fix that only addressed one (Anderson caught it):

Threat A — free-ridingThreat B — cloning
What happensSomeone tails our picks live without payingSomeone distils our projection function from history
NeedsPicks before tipA pile of (features, proj) pairs, at any age
Fixed by a time delay?Yes — a graded pick cannot be tailedNo. A cloner is perfectly happy collecting post-grade data for three months

Delaying publication is the wrong axis for B. The only lever is omission, and it has to be permanent.

What actually separates them: verifying our record needs the pick, line, price and result. It does not need proj. Cloning needs proj — a continuous target is what makes distillation cheap. With it, a few thousand graded picks suffice. Without it, a cloner is fitting a binary side-selection and recovers the decision boundary, not the underlying number, at far worse fidelity for far more data.

So proof and cloning do not require the same field, and the boundary is by VALUE, not by time:

  • Public, permanently: pick, line, side, price taken, closing price, result, CLV, ROI with error bars, void reasons, this ledger. Everything needed to audit us.
  • Product (subscriber-side) if ever triggered: the continuous proj / hit_prob.

Pre-registered decision, made now so it is not made under temptation: if a validated edge appears (2σ, forward-tested, per Claim 4) and we have paying subscribers, per-pick proj/hit_prob become subscriber-visible rather than public. The record stays fully open — every pick, price, result and CLV remains auditable, nothing is ever removed, and the tracker still proves everything it proves today. Additionally, and separately, pending picks may publish post-grade to address Threat A.

Two reasons to under-react, both load-bearing:

  1. We do not publish our input features. A cloner must rebuild recent form, minutes and opponent context from nba_api/ESPN themselves. Anyone who can do that is most of the way to building their own projection — the scarce asset was never the function.
  2. By Claim 4 the projection is worth approximately nothing. Cloning it yields a model that measured efficient against the market.

Status: ACTIVE (policy pre-registered 2026-07-20, not yet triggered).


Known measurement failures we corrected

Kept because a ledger without its own errata is marketing.

DateWhat was wrongEffect
2026-07-12Record graded at a refreshed price while displaying the open priceCredited itself a number it never posted; fixed by watch-then-lock
2026-07-12DFS pick'em prices used as both fair value and bet priceBoard read +0.11% EV; restricted to takeable prices it was −4.19%
2026-07-20MLB graded unlocked picks at the drifted line82% of drifts favoured us; 30 results wrong, 20 fake wins
2026-07-20CLV differenced prices quoted on different linesMLB record CLV read −0.16%; correctly measured it is +1.07%
2026-07-20The public share card computed CLV with its own copy of the mathThe most public artifact disagreed with the audit log it invited people to check
2026-07-20Claim 3's hold table used already-shopped prices as "a book's own hold", and measured the shopped column on locked picks onlyReported 3.56% → 1.75%; real single-book holds are 7–9%. Both errors flattered us
2026-07-20Claim 6 proposed publishing proj post-grade to prevent cloningA time delay stops tailing, not cloning — a cloner collects historical pairs happily. The lever is omission, not delay
2026-07-20Claim 2 cited Pinnacle's prop vig as evidence that no sharp reference existsMeasured: Pinnacle prices both sports at ~7%, mid-field vs US books (6.4–8.9%). The number was right, the inference was refuted — ~7% is just what props cost

Every number below is read from the live boards and written here mechanically — nobody types them in.

The record (play=1, graded, from WNBA 2026-07-08 / MLB 2026-07-12)

sportW-LnROIsignificanceplays needed for 2σ
WNBA382-314696+5.62%+1.4σ~1,356
MLB2096-16263722+8.29%+4.2σ~841

A result under ~2σ is not distinguishable from zero. Where a sport clears that bar we report its record as measured history — never a per-bet +EV claim or a forecast; where it does not, we make no profitability claim at all. Significance is clustered by game-day (see Claim 1).

CLV (diagnostic, not a gate — see Claim 2)

sportavg CLVbeat closenexcluded: line relisted
WNBA+0.26%1661/30933093891 of 4367 graded
MLB+1.27%3387/49524952211 of 11483 graded

Break-even CLV ≈ fair × hold ≈ +2.25 points. A positive CLV below that is still a losing price. Exclusions are picks where the book relisted the line: unmeasurable, not zero.

Is the closing line sharper than the open? (stratified by market × line × side)

sportopen AUCclose AUCgapn
WNBA0.5230.516-0.0062,194
MLB0.5560.562+0.00510,596

A gap near zero means the close carries no more information than the open — nothing is moving these lines toward truth. This is the measurement that demotes CLV to a diagnostic.

The hold we pay, and the two things that lower it

sportmarketbooksmedian bookbest single bookline-shoppedafter lock rulen
—(no per-book ladders stored for this sport)

Two SEPARATE mechanisms, both measured per pick from the per-book ladders at one instant:

  • Line-shopping — median book → best-of-N. What taking the best price buys.
  • Lock discipline — best-of-N → the hold on picks we actually lock. What waiting for a mature market buys on top of shopping.

The honest marginal number is line-shopped vs BEST SINGLE BOOK, not vs the median: a user who simply bets at the sharpest book they already have captures most of the shopping benefit with no tool at all. Quote that gap, not the headline one.

Break-even edge is half the hold. Note the 'after lock rule' column is conditioned on the lock rule itself (locking requires hold ≤ threshold), so it is what we pay — not what shopping achieves.

How we measure accuracy · StatIQ Sports