Field Note 06 / Backtesting
A 100% Win Rate, Without a Proven Trading Edge
Klines logged 15 wins from 15 trades in its strongest paper tier. The number is real. It is also partly guaranteed by the way that tier and its exit are defined. The useful question is not whether 100% appeared in the log, but what that number actually proves.
The result that looked perfect
The newest Klines snapshot covers 83 signals recorded over 19 days, from 19 August to 7 September 2026. This is not a conventional in-sample trading backtest: it is a forward paper log using the real detection timestamp and the reference-price format described in Research Note 06. It is still paper execution, not live trading.
If every candidate produced by the monitor had been entered, the full population would have lost 46.64% in aggregate, with a 39.8% win rate. But “enter everything” is not the trading strategy. Klines includes a Stage-3 gate specifically to reject most candidates before a discretionary trader acts.
That gate separated the sample cleanly. The 33 confirmed signals produced a paper sum of +150.34%. The 41 cancelled signals produced −151.23%, and nine warnings produced −45.75%. Filtering to the 31 ALIVE labels a trader could actually act on produced a 74.2% paper win rate, a +4.95% mean, and a +2.97% median.
The mechanism behaved as designed: it separated the accepted side from the rejected side. That is a real result. It is not yet the same thing as a profitable quantitative trading system.
Why the 100% win rate was partly arithmetic
The headline sits inside the strongest ALIVE tier, internally called STRONG_CLEAN. It closed 15 of 15 paper trades positively, for a 100% win rate and a +143.56% aggregate result.
That number is accurate, but it is weak evidence on its own. The tier is defined partly by a strong upward move in the first minutes, while the exit uses a trailing stop. Those two rules create an arithmetic floor: by the time a candidate qualifies, a non-negative paper outcome is close to guaranteed by construction.
Four of the 15 trades exited exactly at that floor. They contributed nothing beyond the mechanic. Measured above the floor, the added result was a much smaller +1.42 percentage points at the median. That smaller number says more about possible signal quality than the perfect win rate does.
Klines behaved as designed. The mistake would be interpreting correct mechanical behavior as proof of a profitable trading edge.
The gate works—but that is not profit
A filter can be useful even when it is not independently predictive. Here the Stage-3 gate divides a positive paper side from a negative one. For an algorithmic trading system, that is evidence that the classification rule is doing something observable rather than assigning labels at random.
But the cancelled trades stopped being monitored at cancellation. Their full 60-minute paths are missing. The log can show that cancelled candidates were losing when the system rejected them; it cannot show whether they later recovered. There is no measured counterfactual yet, so the economic value of cancellation remains open.
The positive side also has concentration risk. One +43.16% outlier carries a substantial part of the total. Removing it cuts the traded paper sum from +153.59% to +110.43%. The result stays positive, but the mean roughly halves.
The ordinary signal is close to flat
Outside the partly mechanical top tier, the ordinary confirmed signal is not carrying the result. The GOOD_FLOW tier contains 14 trades with a mean of +0.37%, a median of −0.15%, and a 42.9% win rate. On this small sample it is approximately break-even before real execution costs.
This is why backtesting trading strategies requires more than reporting the best-performing bucket. A useful audit asks where the return comes from: a broad repeatable effect, a rule that overlaps with the measured outcome, one unusually large trade, or a fill assumption the live market may never provide.
Paper-positive is not a live edge
Every result in this snapshot is priced at a reference price recorded after Stage-3 confirmation, when part of the qualifying move has already happened. A live order still has to arrive, cross or rest in the book, fill at an available price, and survive spread, slippage, fees, and latency.
This is not a minor footnote. An earlier, more pessimistic fill assumption in the same project flipped a comparable population from a paper gain to a paper loss simply by changing the assumed entry. Until Klines records real fills, the distance between its reference price and an executable price is the decisive unknown.
How to audit a perfect backtest
- Check whether the label contains the outcome.If a tier requires a price move that also creates the measured profit, some success is mechanical.
- Measure beyond the built-in floor.Separate the return guaranteed by the rule from the return added after qualification.
- Remove the largest winner.A strategy that depends on one outlier is a different proposition from one supported by its median trade.
- Inspect the ordinary tier.The less selective signal shows whether the effect is broad or concentrated in a special case.
- Preserve rejected paths.Without the counterfactual, a filter cannot prove how much loss it prevented.
- Replace reference prices with fills.A paper edge becomes a trading claim only after realistic execution and costs are observed.
What Klines has demonstrated
The detector runs. The Stage-3 gate separated a paper-positive accepted side from a negative rejected side. The strongest tier behaved exactly as its rules imply. Those are useful engineering and research findings.
The 100% win rate is not fake. It is simply answering a narrower question than the headline number appears to answer. Klines has produced a promising paper-execution result. Whether that result survives real fills remains open.
Full first-party record: Klines — Paper Snapshot, 7 September 2026 →
