StatIQ Claims Ledger
Every public claim we make, the query that backs it, and the result that would prove it wrong.
This exists so that when someone asks "why do you say that," the answer is a number and a method, not an opinion. Others are free to disagree — but they have to disagree with the data, and they can check it themselves.
Three rules for this document:
- The falsifier is written before the numbers come in, not after. "That metric doesn't apply to us" is a legitimate methodological position and also the exact sentence a losing bettor reaches for. The only thing that separates them is whether the test was specified in advance. If we ever want to drop a claim, we change it here first, in the open.
- The unflattering entries stay. A ledger you would not hand to a skeptic is not a ledger. The entries below where we measured against ourselves are what make the rest credible.
- Numbers are generated, never typed. Every figure in the tables below is read straight off the live boards by a script and written into this document mechanically. Nobody types a number into it, so it cannot quietly fall out of date and it cannot be nudged.
Claim 1 — We make no ROI or +EV claim
What we say: we publish a tracked record, every pick logged before the event, wins and losses both, and we do not claim it proves profitability.
Why: per-bet ROI has a standard deviation near 100%, so a few hundred plays cannot separate a real edge from noise. See the significance column and the "plays needed" column in the table below — those are the honest error bars on our own record.
Falsifier / graduation: we may claim profitability when ROI reaches 2σ from zero on n ≥ 1,500 tracked plays, and not before, regardless of how good the interim numbers look.
Status: ACTIVE. This is the standing constraint on all marketing and social copy.
Claim 2 — CLV is a diagnostic in WNBA/MLB props, not a pass/fail gate
What we say: we publish our closing-line value, and we do not treat negative CLV in these markets as evidence we are wrong.
Why: CLV is a proxy. Its entire validity rests on the premise that the closing line approximates the true price — which requires informed money to move it there. We tested that premise on our own board and it failed:
- Stratified close-vs-open AUC gap ≈ +0.008 (see table). The close predicts outcomes no better than the open. Stratification by market × line × side is mandatory; pooled, this test measures line level and reports a large fake gap.
- Matched line-move test: comparing moved vs un-moved lines within the same market and same open line, movement produced no lift (up-moves −0.5 pts, −0.1σ; down-moves −3.9 pts, −0.6σ).
- What looked like movement mostly wasn't. 184 of 251 MLB "moves" were exactly 1.0 on markets whose only lines are 0.5 and 1.5 — alt-line relisting, not a market moving.
A proxy that fails its own premise does not overrule a direct measurement. ROI is the primary measure. CLV is published as a diagnostic with this caveat attached.
Scope limit — this does not travel. It applies to thin prop markets we have measured. For NBA and NFL sides the close is very likely informative, and CLV remains a gate there.
Falsifier: if the stratified AUC gap in these markets reaches +0.03 or more at n > 5,000, the close is carrying information, and CLV returns to being a gate.
PINNACLE: MEASURED 2026-07-20, and the argument we were making is REFUTED. We had claimed
"Pinnacle prices these props at ~7% vs ~2% on core markets, so it declines to stand behind its number
and no sharp reference exists." We had never pulled it — our production request uses regions
us,us2,us_ex,us_dfs and Pinnacle lives in eu. We went and pulled it directly, and that settles it:
| Pinnacle | US field (same events, same moment) | |
|---|---|---|
| WNBA props | 6.96% (28 legs) | 6.41% – 8.85% across 9 books |
| MLB props | 6.99% (40 legs) | 6.38% (BetOnline) |
The 7% number was right. The inference from it was wrong. Pinnacle sits squarely in the middle of the field — tighter than DraftKings (7.53%), ESPNBet (8.78%) and BetParx (8.85%), looser than BetOnline (6.41%) and Novig (6.44%). ~7% is simply what player props cost across the entire industry; it is not Pinnacle declining to price them. It prices them, on both sports, at market-typical margin.
Which means the hold argument is a dead end in BOTH directions. Margin tells us about a book's pricing model, not about the accuracy of its number. We cannot conclude "no sharp reference exists" from vig, and we equally cannot conclude Pinnacle is sharp from vig.
The real test is still undone: score our graded picks against Pinnacle's de-vigged number and ask whether it predicts outcomes better than our consensus close does. That is the only thing that would confirm or kill Claim 2, and it requires either historical Pinnacle odds (billed 10×) or forward collection from now. Until it runs, Claim 2 rests on the stratified AUC test and the matched-move test only — and the Pinnacle line must not be cited as supporting evidence.
Status: ACTIVE (established 2026-07-20).
Claim 3 — Line-shopping and lock discipline cut the hold we pay by ~64%
What we say: taking the best available price across books materially lowers the edge required to break even, and this is true whether or not our projections are any good.
Why: it is mechanical, not predictive. This is the most defensible thing we do — it requires no model to be right, and it is verifiable on any single pick. Measured per pick from the per-book ladders at a single instant (MLB, n=1,011):
median single book 7.29%
best single book 5.95%
line-shopped 4.29% <- shopping buys 3.00 pts vs median, 1.67 pts vs best book
after our lock rule 2.65% <- waiting for a mature market buys another 1.64 pts
Two distinct mechanisms — shopping and lock discipline — each separately measurable. Combined, 7.29% → 2.65% is a 64% reduction in the hold, i.e. break-even edge falls from 3.6 pts to 1.3 pts.
Honest caveats, stated up front:
- Quote the vs-BEST-BOOK number (1.67 pts), not the vs-median one (3.00). A user who simply bets at the sharpest book they already have captures most of the benefit with no tool at all. The median-book comparison flatters us by assuming the user bets randomly.
- The shopped figure is a synthetic cross-book pair — what shopping buys, not a two-way market anyone could bet both sides of.
- The "after lock rule" figure is conditioned on the lock rule itself (locking requires hold below a threshold). It is what we pay; it is not an achievement of shopping.
- We have no WNBA numbers here. WNBA stores per-book prices for our side only, so a genuine single-book hold cannot be computed from stored data. This claim is currently MLB-only.
Correction on record (2026-07-20): the first version of this table reported "3.56% → 1.75%" and was
wrong twice, both times in our favour. It used close_over/close_under as "a book's hold" when those
are already best-of-N across books, and it measured the shopped column on locked picks only — selecting
on the very threshold that defines locking. Real single-book holds are 7–9%, not 3.5%.
Falsifier: if line-shopped hold stops beating the BEST SINGLE BOOK by at least 1 point — because we lose book access, or because the books we can reach converge — this claim dies with it. Tracked in the generated table below, which reports both comparisons every run.
Status: ACTIVE.
Claim 4 — The prop markets we have measured are efficient, and we say so
What we say: WNBA and MLB player props measured efficient against our model. We do not claim to have found a persistent edge in them.
Why: this is the entry we would most like to delete, which is exactly why it stays. Repeatedly through 2026, apparent edges found at n < 150 dissolved at larger n: the blowout-under thread (full season, 302 games, −5.31%), "unders survive" (noise, +1.2σ, calibration sign reversed), the WNBA totals regime story (+1.9σ in one season, −1.3σ in the other, +0.1σ combined). The model is frozen against tuning until at least 2026-07-25 for this reason — we are collecting data, not fitting it.
Falsifier / graduation: an edge is real when it clears 2σ on a pre-registered, forward-tested sample — never on the sample that suggested it.
Status: ACTIVE.
Claim 6 — We publish what PROVES us, not everything that BUILDS us
What we say: our record, our grading, our method and our mistakes are open and checkable. Our per-pick model inputs are not a permanent public commitment.
Why this needs deciding in advance. Everything we publish is public JSON: 8,323 MLB and 1,953 WNBA
picks, each carrying proj, hit_prob, ev, edge, the line, and the graded actual. That is a
labelled dataset of our model's outputs. Anyone competent with an AI assistant can scrape it and train a
surrogate that reproduces our projections to within noise — no code, no injury feed, no minutes model
required. A subscriber could pay for one month and walk out with a working clone. This is not a
hypothetical; it is a weekend of work for a technical person.
The distinction that resolves it:
- Numbers that PROVE us — every pick logged before tip, the result, the price we took, the closing price, CLV, ROI with error bars, the void/DNP reasons, this ledger and its errata. These stay open. They are the entire basis of the claim that we are honest, and hiding any of them would forfeit the one position nobody else in the category occupies.
- Numbers that BUILD us — the per-pick projection and hit probability on pending picks, i.e. the live model output. These are open today by default, not by decision, and that default is what a cloner needs most.
Current exposure is genuinely zero, on two independent grounds: no one has signed up yet, and our own measurements say the projection adds ~nothing beyond the market ([[Claim 4]]) — so a clone of it reproduces a model that measured efficient. Both of those change the moment either one changes.
TWO DIFFERENT THREATS. The first version of this claim conflated them and proposed a fix that only addressed one (Anderson caught it):
| Threat A — free-riding | Threat B — cloning | |
|---|---|---|
| What happens | Someone tails our picks live without paying | Someone distils our projection function from history |
| Needs | Picks before tip | A pile of (features, proj) pairs, at any age |
| Fixed by a time delay? | Yes — a graded pick cannot be tailed | No. A cloner is perfectly happy collecting post-grade data for three months |
Delaying publication is the wrong axis for B. The only lever is omission, and it has to be permanent.
What actually separates them: verifying our record needs the pick, line, price and result. It does
not need proj. Cloning needs proj — a continuous target is what makes distillation cheap.
With it, a few thousand graded picks suffice. Without it, a cloner is fitting a binary side-selection
and recovers the decision boundary, not the underlying number, at far worse fidelity for far more data.
So proof and cloning do not require the same field, and the boundary is by VALUE, not by time:
- Public, permanently: pick, line, side, price taken, closing price, result, CLV, ROI with error bars, void reasons, this ledger. Everything needed to audit us.
- Product (subscriber-side) if ever triggered: the continuous
proj/hit_prob.
Pre-registered decision, made now so it is not made under temptation: if a validated edge appears
(2σ, forward-tested, per Claim 4) and we have paying subscribers, per-pick proj/hit_prob become
subscriber-visible rather than public. The record stays fully open — every pick, price, result and
CLV remains auditable, nothing is ever removed, and the tracker still proves everything it proves today.
Additionally, and separately, pending picks may publish post-grade to address Threat A.
Two reasons to under-react, both load-bearing:
- We do not publish our input features. A cloner must rebuild recent form, minutes and opponent
context from
nba_api/ESPN themselves. Anyone who can do that is most of the way to building their own projection — the scarce asset was never the function. - By Claim 4 the projection is worth approximately nothing. Cloning it yields a model that measured efficient against the market.
Status: ACTIVE (policy pre-registered 2026-07-20, not yet triggered).
Known measurement failures we corrected
Kept because a ledger without its own errata is marketing.
| Date | What was wrong | Effect |
|---|---|---|
| 2026-07-12 | Record graded at a refreshed price while displaying the open price | Credited itself a number it never posted; fixed by watch-then-lock |
| 2026-07-12 | DFS pick'em prices used as both fair value and bet price | Board read +0.11% EV; restricted to takeable prices it was −4.19% |
| 2026-07-20 | MLB graded unlocked picks at the drifted line | 82% of drifts favoured us; 30 results wrong, 20 fake wins |
| 2026-07-20 | CLV differenced prices quoted on different lines | MLB record CLV read −0.16%; correctly measured it is +1.07% |
| 2026-07-20 | The public share card computed CLV with its own copy of the math | The most public artifact disagreed with the audit log it invited people to check |
| 2026-07-20 | Claim 3's hold table used already-shopped prices as "a book's own hold", and measured the shopped column on locked picks only | Reported 3.56% → 1.75%; real single-book holds are 7–9%. Both errors flattered us |
| 2026-07-20 | Claim 6 proposed publishing proj post-grade to prevent cloning | A time delay stops tailing, not cloning — a cloner collects historical pairs happily. The lever is omission, not delay |
| 2026-07-20 | Claim 2 cited Pinnacle's prop vig as evidence that no sharp reference exists | Measured: Pinnacle prices both sports at ~7%, mid-field vs US books (6.4–8.9%). The number was right, the inference was refuted — ~7% is just what props cost |
Every number below is read from the live boards and written here mechanically — nobody types them in.
The record (play=1, graded, from WNBA 2026-07-08 / MLB 2026-07-12)
| sport | W-L | n | ROI | significance | plays needed for 2σ |
|---|---|---|---|---|---|
| WNBA | 42-52 | 94 | -4.37% | -0.4σ | ~2,406 |
| MLB | 111-97 | 208 | +5.12% | +0.7σ | ~1,579 |
A result under ~2σ is not distinguishable from zero. We make no ROI or +EV claim at these sample sizes, regardless of sign.
CLV (diagnostic, not a gate — see Claim 2)
| sport | avg CLV | beat close | n | excluded: line relisted |
|---|---|---|---|---|
| WNBA | +0.50% | 147/249 | 249 | 63 of 1689 graded |
| MLB | +1.02% | 531/810 | 810 | 83 of 6360 graded |
Break-even CLV ≈ fair × hold ≈ +2.25 points. A positive CLV below that is still a losing price. Exclusions are picks where the book relisted the line: unmeasurable, not zero.
Is the closing line sharper than the open? (stratified by market × line × side)
| sport | open AUC | close AUC | gap | n |
|---|---|---|---|---|
| WNBA | 0.581 | 0.584 | +0.003 | 120 |
| MLB | 0.569 | 0.574 | +0.005 | 3,028 |
A gap near zero means the close carries no more information than the open — nothing is moving these lines toward truth. This is the measurement that demotes CLV to a diagnostic.
⚠️ Underpowered: WNBA. Only strata with ≥40 graded picks are included, so a small n here means most of the board could not be tested at all — not that the test came back clean. The claim rests on the larger sample; this row is not independent confirmation.
The hold we pay, and the two things that lower it
| sport | market | books | median book | best single book | line-shopped | after lock rule | n |
|---|---|---|---|---|---|---|---|
| MLB | batter_hits | 8 | 7.02% | 5.91% | 4.43% | 1.77% | 270 |
| MLB | batter_home_runs | 3 | 8.89% | 7.49% | 5.79% | — | 218 |
| MLB | batter_rbis | 6 | 7.26% | 5.95% | 3.23% | — | 265 |
| MLB | batter_total_bases | 8 | 7.24% | 5.95% | 4.14% | 3.06% | 258 |
Two SEPARATE mechanisms, both measured per pick from the per-book ladders at one instant:
- Line-shopping — median book → best-of-N. What taking the best price buys.
- Lock discipline — best-of-N → the hold on picks we actually lock. What waiting for a mature market buys on top of shopping.
The honest marginal number is line-shopped vs BEST SINGLE BOOK, not vs the median: a user who simply bets at the sharpest book they already have captures most of the shopping benefit with no tool at all. Quote that gap, not the headline one.
Break-even edge is half the hold. Note the 'after lock rule' column is conditioned on the lock rule itself (locking requires hold ≤ threshold), so it is what we pay — not what shopping achieves.