the auditable spec, release r2
Method: built to be attacked.
Everything below is reproducible from raw files; a skeptical quant is the intended reader. The repository version of this document contains the full parameter table and the deviations log.
Data
CFTC Commitments of Traders, full official history via the CFTC’s public API: legacy report (futures-only and futures+options, 1986→, including trader counts and concentration ratios), disaggregated report (2006→, managed money / swap dealers / producers), and Traders in Financial Futures. Raw copies are stored immutable before any transformation; duplicate rows are reconciled and logged; positioning data is never forward-filled.
Prices are continuous front-month series (proxy, roll methodology not under our control; flagged on every payload). Term-structure conditions are parked until a back-adjusted source is added, rather than computed from unfit data.
Publication lag · the zero-lookahead rule
COT data is as-of Tuesday, released Friday ~15:30 ET. Every forward return, threshold and drawdown figure is measured from the first session OPEN after release, the first price a real trader could have had. An automated test shifts all COT data forward a week and verifies signals move accordingly.
Extreme definitions & de-clustering
Seven fixed definitions (COT index ≥95/≤5, net %OI percentile ≥95/≤5, 156-week z-score ±2, flush-at-extreme, extreme + price divergence, extreme + 4-week price failure, commercial capitulation), all on trailing windows only, parameters were fixed before the confirmation period was ever read.
Consecutive weeks at an extreme are ONE episode: first crossing counts, a new event requires 13+ weeks and a genuine exit from the zone in between.
Outcomes, baselines, corrections
A reversal is a counter-crowd move of 0.75× the market’s own volatility scaled to each horizon (1–26 weeks), from the entry-eligible open. Every conditional rate is shown next to the unconditional base rate of the same move.
Inference uses placebo resampling (2,000 min-gap pseudo-event draws per cell) and Benjamini–Hochberg FDR at q=0.10 across the full grid of ~3,700 cells. Train period 2006–2018, confirmation 2019→. Cells with fewer than 8 train episodes are labeled insufficient and never presented as edges.
Every return, bar and drawdown on this site is the SIMPLE return on the position: what the money did. The study itself measures in natural log, because that is the space where returns add up across windows and where the volatility bar is defined, and the frozen record stays there. Since 2026-07-28 the conversion happens once, where the data is loaded, and it is direction-aware: a fade LONG earns exp(x)−1 of the move and a fade SHORT earns 1−exp(−x), which are not the same function. A consequence log space was hiding: a short can lose more than the whole position, so two cases on this site now read past −100%, and that is the arithmetic, not a typo.
The release r1 finding
Zero of 930 testable market/definition/horizon cells and zero of 232 condition cells survived correction plus out-of-sample confirmation. Pooled base rates run 33–39%. A positioning extreme is a risk state, not a countdown, and every number behind that sentence is on file, per market, on this site.
Historical results are frozen per research release; the weekly job updates current-state readings only. Every published statistic carries its sample size, baseline, confidence interval, correction status and out-of-sample status.
Release r2: what we corrected, and what it did not change
An instrument audit of all 44 markets, checked against the contract names in the CFTC filing itself, found that 40 markets priced exactly the futures contract their report covers. Four did not. BITCOIN and ETHER were priced on SPOT while their COT rows describe CME futures. The graded record you read for those two was rebuilt on BTC=F and ETH=F, which pre-date their own COT series, and the swap changed no outcome. The other two are VIX and USDX, below.
That correction was only half applied for a while, and the half that was missing was the half you look at. The graded record ran on the futures; the chart, the current-state reading and the open-trade card on those two pages ran on spot, because the feed that refreshes this site every week carries spot and overwrote the futures series every Saturday. Since 2026-07-28 all of it runs on BTC=F and ETH=F, and the weekly job now refuses to serve any market from an instrument this project does not configure for it, loudly, instead of reverting to one in silence.
Finishing it left the frozen record alone and moved one live number, and we would rather point at that number than let you find it. The frozen record did not move: every graded episode outcome for bitcoin and ether is exactly what it was before, because it was already measured on the futures. The nowcast scorecard did move. That is the ‘% right’ on those two pages, and it is scored against the daily price series, which for these two markets was the spot one. Rescoring it on the futures takes bitcoin from 65% to 58% and ether from 75% to 67%, and turns three of bitcoin’s last 26 weekly marks and two of ether’s from a hit into a miss. No part of the method changed. We had been marking those two forecasts against a price that was not the one the report describes, and the lower number is the honest one.
For the record of how far apart the two series were: across the overlap they sat about 0.7% apart at the median, and more than 3% apart on 5% of days for bitcoin and 7% for ether. Close enough not to flip a volatility-scaled reversal threshold, and further apart than a basis you could ignore without saying so.
VIX is different, and worse. Its COT row describes CBOE VIX FUTURES. The only series available to us is the VIX INDEX, which is not tradeable and whose returns lack the roll yield that dominates what a futures position actually earns. Every outcome we could publish there would be measured on the wrong instrument, so from r2 on we publish none: VIX keeps its positioning (real CFTC data) and its chart, and contributes zero episodes to the pooled inference. Its 35 episodes left the sample. The headline moved from 39.6% to 39.7%, and the edge over baseline from +1.8 to +1.5 points. The correction bought integrity, not an edge.
USDX is the fourth, and we kept it. Its COT row describes ICE Dollar Index futures (DX); we price it on the dollar index itself (DX-Y.NYB), because no usable DX futures series is available to us on either provider. We kept this one where we dropped VIX because the two cases are not alike: DX futures track the index with a small interest-rate basis and their returns are near-identical, whereas VIX futures earn a roll yield that dominates what the position actually makes, so the index says almost nothing about it. USDX therefore keeps its outcomes, and this paragraph is the disclosure that decision always required. If you want the record measured on the contract rather than the index, USDX is the one market here where you do not have it, and the 29 graded extremes on that page are measured on the index. Every market's price instrument is named in the price_instrument column of markets.csv in the research pack.
Two price defects are now rejected rather than repaired, and logged: CRUDE_WTI’s negative settle on 2020-04-20 (real history, but a negative price has no logarithm) and a JPYUSD decimal slip on 2001-12-17. An earlier version of that screen also deleted the 2018-02-05 volmageddon crash and the 2010-05-06 flash crash, real history that has the same shape as a bad tick; the rule now only rejects prints that sit a clean factor of ten from their neighbours. We never patch, interpolate or smooth a price: an invented number is worse than a missing one.
What we still cannot rule out: our roll-artefact test only catches splices that reverse, so a roll into a persistently higher contract (contango) would be missed. The 1.5% of episode windows we flagged is a floor, not a ceiling. Closing that gap needs a back-adjusted reference series we do not yet own, and we would rather say so than imply the question is settled.
On trial: popular entries, and our own protocol
The honest way to sell a map is to also publish the routes that failed. Both studies below were pre-registered: rules fixed before running, no parameter search, train 2006–2018, confirmation 2019 to today, daily data, release-aligned (no lookahead). Scripts ship in the repository (research/sd_cot.py, research/protocol_test.py).
Supply/demand-zone retests (1,075 events, 44 markets) · mean 10-session ATR-scaled return:
| variant | train | confirm | n |
|---|---|---|---|
| zones standalone | −0.17 | −0.14 | 1,075 |
| zones only inside a COT extreme (fade-aligned) | −0.38 | −0.78 (t −2.4) | 56 |
| plus the fingerprint filter | −0.37 | −1.06 | 28 |
Supply/demand-zone retests · 1,075 events across 44 markets.
Verdict: supply/demand retests were negative in both eras, standalone and with COT. Over the confirmation window the same dates carried a +0.31R unconditional drift: the zone logic did worse than doing nothing. The verdict applies to this registered zone definition; we did not try others until one “worked”.
Our own protocol, mechanized (extreme, fingerprint gate, first counter-close entry, median-shakeout stop, 13-week clock · 1,050 trades):
| variant | train | confirm | stopped out |
|---|---|---|---|
| extreme + trigger, no gate | −0.06 | −0.16 | 81–83% |
| plus the fingerprint gate (≥2 of 3) | +0.20 | −0.04 | 78–82% |
Our own protocol, mechanized · 1,050 trades.
Verdict: null against the registered bar. Even our own rules, mechanized, carry no entry edge, and we say so. Two measurements survived: the fingerprint gate improved results by the same sign in both eras (+0.50R train, +0.22R confirmation vs the trades that failed the gate; +0.49R/+0.21R after a within-market control, so it is not market mix; the out-of-sample gap is modest and not statistically significant), which is why the product treats it as a priority map and still not as a signal; and a tight stop (the larger of the market’s median shakeout and one ATR, the ATR floor binding in 60% of trades) got stopped out of 81% of everything, which is why the risk pages say to size from the worst case, not the middle.
Seasonality got the same treatment. We tested a “second half of the year” turn condition across all 44 markets and every asset group; nothing survived multiple-testing correction (best raw p = 0.05, no FDR pass). It is on the record with the rest of the nulls, not in any claim.
Every setup we tested and killed → · the graded record itself →