the setups we tested and did not sell
Everything here failed.
We published it anyway.
Every setup below is one somebody is selling you right now. We fixed the rules in writing before running them, tested each on the full record, and kept the results whichever way they came out. They came out flat or negative · including our own idea, which is the one that cost us the most.
pre-registered · train 2006–2018, confirmation 2019 onward · no parameter search · every script ships in the repository
Why this page exists. Anyone can show you a backtest that worked. The only thing that tells you whether a research process is honest is what it does with the tests that did not. This is ours, in full. It is also the reason we do not sell a signal: we went looking for one, five separate ways, and the record says it is not there at weekly COT granularity.
Blind fading of the extreme
1,234 tradesThe setup every COT account sells: positioning hits an extreme, so bet against the crowd.
| variant | train | confirmation |
|---|---|---|
| fade the crowd at the extreme | −0.033R | −0.036R |
| the mirror: follow the crowd instead | — | −0.024R pooled |
metric: expectancy per trade, in R (risk units), 0.10% costs
50% hit rate, profit factor 0.91, negative in 17 of 21 years, and statistically identical in both eras. The point estimate sits inside its own noise band, so we do not claim it proves a loss. What it does show is that there is no directional edge at this granularity, in either direction, and costs eat a coin flip. Fading and following are both flat.
research/strategy.py
Waiting for confirmation before entering
three pre-registered levels, all publishedThe confluence doctrine: do not catch the falling knife, wait until the reversal confirms.
| variant | train | confirmation |
|---|---|---|
| V1 · price break + first spec cut, both in | −0.046R | profit factor 0.88 |
| V2 · break + a 0.25σ√13 counter-move | −0.045R | −0.071R |
| V3 · full reversal unit in, ride the new trend | −0.058R | −0.109R |
metric: expectancy per trade, in R
This is the most useful result on the page, and it inverts what every confluence course teaches: the more confirmation you demand, the WORSE it gets, monotonically. Each added layer costs roughly 0.02R. Later entry surrenders move while stop noise stays constant. Confirmation buys comfort, not edge, and the data prices the comfort.
research/strategy2.py
Supply/demand zone retests
1,075 zone events, 44 marketsThe most-taught discretionary entry there is, tested standalone and then inside a COT extreme.
| variant | train | confirmation |
|---|---|---|
| zones standalone | −0.17 | −0.14 |
| zones only inside a COT extreme (n=56) | −0.38 | −0.78 |
| plus the fingerprint filter (n=28) | −0.37 | −1.06 |
metric: mean 10-session ATR-scaled return · the small-n rows below the first are shown because we registered them, not because they conclude anything
Negative in both eras standalone, across 1,075 events. Over the same confirmation window the identical entry dates carried a +0.31R unconditional drift, so the zone logic did worse than doing nothing at all. The COT extreme does not rescue the zone. The two variants measured inside a COT extreme are n=56 and n=28: too few to conclude from, and shown only because we registered them in advance.
research/sd_cot.py
Our own protocol, mechanised
1,050 tradesThe obvious next move: only take the trade when the conditions that historically accompanied turns are present.
| variant | train | confirmation |
|---|---|---|
| extreme + trigger, no condition gate | −0.06 | −0.16 |
| plus the fingerprint gate (≥2 of 3 conditions) | +0.20 | −0.04 |
metric: mean R per trade, stop = max(median shakeout, 1 ATR)
Read the gated variant twice, because it is the whole argument of this site. Gating on the historically reliable conditions produced +0.20R in the years the rule had already seen, and nothing out of sample. A positive result in-sample that vanishes out-of-sample is the textbook signature of overfitting, and it is the exact pattern our methodology exists to catch. It also stopped out about 81% of its trades. We tested our own idea and it failed, so we are not selling it.
research/protocol_test.py
Calendar seasonality, alone and with COT
8,897 eligible market-months since 2006Trade a market's seasonal tendency: buy the months it has historically risen, sell the ones it has historically fallen · and stack it on a COT extreme for confirmation.
| variant | train | confirmation |
|---|---|---|
| train 2006–2018 (n=5,102) | +0.34% timed | +0.36% simply held |
| confirm 2019→ (n=3,795) | +0.69% timed | +0.87% simply held |
| at a COT extreme, confirm | gap +1.32% | but train gap fails |
metric: mean one-month return, and the same return benchmarked against simply holding
Seasonality looks profitable and is not. The timed rule is long about 61% of months and these markets drifted up, so the positive figures are drift wearing a calendar: benchmarked against simply holding every eligible month long, seasonal timing LOSES in both windows (−0.02 pts train, −0.19 pts confirm). Stacked on a COT extreme it looks better · fades did 1.32 points better when the season agreed · but that gap fails in training and clears out of sample by a hair (interval lower bound +0.02%), which is one window of noise, not a finding. We also disclose a defect in our own bar: the adoption threshold was mis-scaled by a factor of twenty and the bootstrap interval it called for was never computed. Both were fixed after we had seen the numbers, so this study is exploratory rather than pre-registered, and we say so instead of presenting a repaired bar as the original.
research/seasonality.py
Larry Williams' commercials rule
1,657 commercial-extreme eventsFollow the commercials: when their positioning is at an extreme of its 3-year range, trade with them.
| variant | train | confirmation |
|---|---|---|
| train 2006–2018 (n=1,065) | 50.1% at 8w | 51.8% at 13w |
| confirm 2019→ (n=592) | 53.2% at 8w | 50.2% at 13w |
metric: hit rate vs the same-direction baseline drift (49.9%)
A coin flip, and we are explicitly NOT claiming a kill. The confirm-era 8-week reading is nominally above the 49.9% baseline, the edges do not point consistently across era and horizon, and no significance test was registered for this study, so none is claimed. Two honest caveats: average signed returns turn negative in the confirmation era while the hit rate rises, and Williams himself presented this as a condition to be paired with price triggers, never as a standalone timing rule. We tested it as a timing rule, which is not quite what he sold.
research/williams.py
So why pay us, if we debunk everything?
Look at what those studies actually tested. Every one of them mechanized the entry: go here, exit there, no judgement in between. That is the thing being sold to you everywhere, and that is the thing that failed. Not one of them tested your read of the tape, your sizing, or the weeks you decide to sit out. Your edge was never on trial here. We cannot disprove it, and we are not trying to sell you a replacement for it.
What we sell is the ground underneath it. Where the crowd actually sits today, not in Tuesday’s stale report. How far past turns in this market ran against the trade before they worked. What the worst case on file says about size. Which of the 44 markets are even worth your attention this week. All of it counted, none of it forecast. You still make the call · you just make it with the file open.
One number we have never seen a signal seller publish: the protocol above stopped out about 81% of its trades on the only stop rule we tested. Everyone sells a stop. Ask them for that figure.
The Friday Letter
Every Friday 21:30 CET the new report lands and we regrade all 1,206 cases overnight. Confirm and your crowd-map snapshot arrives right away. Then every Friday: what moved, where the crowd is stretched, and which fades are running in public. And the moment we open one, you get the ping. Free.
Every Friday: where the crowd is stretched, what is running, what it cost. Free.
Confirmation by e-mail, one-click unsubscribe in every issue. Your address is never shared.