Roger Mendoza

The 80% win rate that wasn't

A backtest dashboard showed squeeze signals winning 80% of the time. Two lines of code explained why — an outcome-defined label and exits at the peak price.

There's a rule I try to follow with backtests: the better the number, the harder I look at the code.

The 3WT scanner had a "lab" dataset behind one of its dashboards. It split historical signals by whether they were squeezes and whether momentum was on, then simulated LEAP options trades. The headline was hard to ignore: 80% wins and +1,609% total return for squeeze-plus-momentum signals, with a maximum drawdown under 2%.

Meanwhile, the plain stock-level backtest of the same scanner — covered in the previous note — said 44.8% wins and a profit factor of 1.11. Both can't describe the same market. So I read the code that generated the lab numbers.

Not investment advice — this is about how a simulation misled its own author.

Problem 1: the label was decided by the outcome

The lab grouped trades into "squeeze" and "not squeeze". Here's the definition, simplified:

squeeze = breakout or max_gain_pct >= 20

breakout was checked on bars after the entry, and max_gain_pct is the biggest rise the stock achieved after the entry. In other words, a trade was labelled a squeeze if it went up a lot.

Then the dashboard reported that squeezes win 79% of the time. Of course they do — winning is part of the definition. This is circular reasoning wearing a lab coat, and it's surprisingly easy to write: the variable name sounds like a property of the setup, while the code measures a property of the result.

Problem 2: exits at the top

For any trade not stopped out, the options model credited the gain like this:

upside = max(max_gain_pct, 0)          # the highest high in the window
option_gain = upside * 0.5 * 2.0 - decay

max_gain_pct is the peak — the single best price the stock touched over the holding window. No real trader sells at the exact top of a 60-day window, because no real trader knows where it is until later. Crediting it is look-ahead bias, and it inflates every held trade.

The time-decay charge didn't help either: 0.003 percentage points per day, which is effectively zero for an option.

Win rate: lab options model vs stock-level backtestLab: squeeze label ✓79.3%Lab: all signals55.6%Stock backtest: all trades44.8%
The lab's 'squeeze' group was labelled using the outcome itself, and held trades were credited at their peak price. The honest stock-level number is the bottom bar.

What the honest number is

The stock-level backtest exits on its rules — stop, target or time — at prices that were actually knowable on the day. Its 44.8% win rate and 1.11 profit factor are the numbers I'd put my name to. The lab's 79–80% describes the simulation, not the market.

Researchers have a name for this family of mistakes. Bailey, Borwein, López de Prado and Zhu called it backtest overfitting: the more freedom a simulation has to choose its labels and exits after the fact, the better it looks and the less it means.

The checklist I use now

  1. Is every input knowable at the time of the decision? Labels, filters and exit prices must use only data from before the trade.
  2. Does a "group" variable secretly measure the result? If the name describes a setup, check that the code does too.
  3. Where does each exit price come from? If it's a max or a min over the future, it's fiction.
  4. Do two backtests of the same idea disagree? Then at least one is wrong, and the more flattering one is the usual suspect.

Finding this felt bad for about ten minutes and good ever since. A number that would have been embarrassing in public was caught by reading forty lines of my own code. That's the cheapest a mistake like this will ever be.

Some links in these notes are affiliate links. If you buy through one, I may earn a commission at no extra cost to you. I only link to tools I use or would recommend.