The EMA 9/20 Pullback Loses Money on ES, and Our First Two Runs Hid It
We re-ran the EMA 9/20 pullback on twelve and a half years of our own ES hourly data: 12,130 trades, 32.8% wins at a 2:1 target, −$37.90 per trade, −$459,730.54 in total, profit factor 0.878. Getting there took three runs. The first filled at the EMA inside the signal bar and printed +$2,290,997; the second scored same-bar stop-and-target as a win and printed −$301,387. The lookahead was worth $2,750,727 and the famous same-bar artifact only touched 1.4% of trades.
The EMA 9/20 pullback on twelve and a half years of ES hourly bars wins 32.8% of its trades at a 2:1 target. Across 12,130 trades that is −$37.90 each, and −$459,730.54 in total. It is the same verdict our NQ study reached in 2026, now on a longer archive and three and a half times the trades.
The reason we are publishing it is not the verdict. It is that we broke this backtest twice before it told us that, and both breaks are the ones sold strategies are built on.
The rule
- ES 1-hour bars, EMA 9 and EMA 20 on the close.
- Long regime when EMA9 > EMA20. Short regime mirrored.
- Signal when the bar trades down into the EMA20 and closes above it. Shorts mirrored.
- Stop one ATR(14) beyond the entry. Target 2× the risk.
- Held at most 59 bars. Whatever is still open then exits at that bar’s close.
- One contract, no filters, no re-entry rules.
First break: we filled inside the signal bar
Our first run entered at the EMA price during the signal bar itself. That is the bar whose close is what tells you a pullback happened. You are buying at a price you only know was a pullback because you already watched the bar finish above it.
It produced this:
| First run, lookahead fill | |
|---|---|
| Trades | 12,130 |
| Win rate | 50.1% |
| Average per trade | +$188.87 |
| Total, one contract | +$2,290,997 |
| Profit factor | 1.821 |
| Max drawdown | $38,689 |
| Sharpe | 6.84 |
| t | 24.38 |
A two-EMA rule with a 6.84 Sharpe and a $38,689 worst drawdown on a $2.3 million curve. If a vendor showed us that, we would assume the fill. We wrote it ourselves.
Moving the entry to the next bar’s open and changing nothing else — same signals, same trades, same stops and targets — moved the result by $2,750,727. The exit mix moved with it. The lookahead version booked 6,071 targets against 5,854 stops. The honest version books 3,979 against 7,983.
Second break: the artifact everyone warns about, and it is the small one
The second run resolved bars that contained both the stop and the target as wins. This is the classic backtest artifact, the one our NQ article spent a section on, and it is genuinely how most naive engines behave.
On ES hourly bars with an ATR stop it touched 166 of 12,130 trades, 1.4%. Scoring them as wins instead of losses moves the result to −$24.85 per trade. In total it is worth $158,344. That is real money, and it is still small next to the fill.
That ordering is worth stating plainly, because it inverts the usual advice: on this timeframe the famous same-bar artifact is the lesser of the two errors. Both convictions are visible in one picture. The green line is the same signals with the fill peeking inside the bar. The two lines below it are the same signals scored two ways once the fill is honest.

The honest result
| ES 1-hour, 2014–2026 | |
|---|---|
| Bars tested | 74,189 |
| Trades | 12,130 |
| Win rate at 2:1 | 32.8% |
| Average per trade | −$37.90 |
| Median trade | −$220.57 |
| Total, one contract | −$459,730.54 |
| Profit factor | 0.878 |
| Expectancy | −0.018R |
| Max drawdown | $512,497 |
| Sharpe | −1.51 |
| t | −5.37 |
| Exits | 7,983 stops, 3,979 targets, 166 bars holding both, 2 timeouts |
| Best / worst trade | +$6,675.86 / −$5,312.54 |
Note what −0.018R means. The rule is not wrong about direction in any dramatic way. It is a coin that lands slightly on the wrong side of its own break-even and pays $29.50 to flip.
Every reward multiple lands under its own line
This is the exhibit we would show first if we could only show one. Each target level has an arithmetic win rate it must clear before costs. We measured all four.

| Target | Break-even | Measured | Per trade | Total | PF |
|---|---|---|---|---|---|
| 1R | 50.0% | 47.3% | −$43.37 | −$526,111.79 | 0.818 |
| 1.5R | 40.0% | 38.5% | −$41.20 | −$499,752.86 | 0.854 |
| 2R | 33.3% | 32.8% | −$37.90 | −$459,730.54 | 0.878 |
| 3R | 25.0% | 24.5% | −$45.10 | −$547,052.86 | 0.871 |
Four configurations, four misses, all short by between 0.5 and 2.7 points. That pattern is the control. A rule with an edge beats its break-even somewhere in a sweep like this and the profitable region tells you where the edge lives. A rule with no edge tracks the line and falls under it by roughly the cost, which is exactly what these four rows do.
The sweep also answers the reward-multiple argument that usually follows a losing test. Widening the target does not help: 3R has the worst per-trade number in the set at −$45.10.
The splits
We cut the sample two ways.
| Trades | Win rate | Per trade | Total | PF | t | |
|---|---|---|---|---|---|---|
| Longs | 6,598 | 34.2% | −$16.65 | −$109,837.43 | 0.94 | −1.88 |
| Shorts | 5,532 | 31.1% | −$63.25 | −$349,893.11 | 0.819 | −5.61 |
| Holdout 2024–2026 | 2,586 | 34.8% | −$22.55 | −$58,303.96 | 0.951 | −1.05 |
Longs lose less than shorts, by $46.60 a trade. On an index that spent most of twelve years rising, a rule that is flat-to-negative on the long side has nothing left to attribute the loss to. The short side at −$63.25 is the same trade run against the drift.
The last two and a half years, held out, tell the same story. The loss is milder at −$22.55 a trade, and it is still a loss across 2,586 trades.
Why the extra five years matter less here than usual
This series exists because our ES archive runs twelve and a half years and our NQ archive runs seven. On the RSI-2 study those five extra years changed the headline — they contained losing calendar years the Nasdaq window could not show.
Here they do not change the verdict, and that is worth saying rather than dressing up. The NQ test found 31% wins at 2:1 across 3,458 trades. The ES test finds 32.8%, on a different index, a different tick value and a window that starts five years earlier. What the extra years bought is sample size. 12,130 trades put t at −5.37, which is enough to call this dead rather than unproven.
What we changed in our own book
We changed the harness, not the book. We were never trading this rule.
- Signal-bar fills are gone from the engine. Any entry priced from a level inside the confirming bar now fills at the next bar’s open. The $2,750,727 gap is why.
- We report both conventions on every intraday study from now on. Same-bar resolution is not the big error, and at 1.4% of trades here it was not close. We publish the pair anyway, because it is cheaper than arguing about it.
- Every fixed-R test gets its break-even row. The four-line sweep above took minutes and made the result unarguable.
The original Nasdaq test is here: the EMA 9/20 pullback on NQ. If you want to re-run this on your own engine, the inputs are ours and they are for sale. ES ticks back to 2014 with the real aggressor side on every print, in the historical data packages.
Methodology: ES 1-hour bars built from our own tick archive, 74,189 bars from 2 January 2014 to 10 September 2026, New York time, all sessions. EMA 9 and EMA 20 on closes; entry on the bar after the pullback bar closes back above the EMA20, at that bar’s open; stop one ATR(14) beyond entry; target 2× risk unless stated; positions held at most 59 bars and closed at that bar’s close if neither level is reached. Bars containing both stop and target are scored as stops. One contract, $4.50 commission plus two ticks slippage per round trip, $29.50 on ES. The NQ comparison figure comes from our 2026 NQ hourly study, 3,458 trades over 2019–2026.
Frequently asked questions
Does the EMA 9/20 pullback work on ES?
No. Over 74,189 hourly bars from January 2014 to September 2026 the rule took 12,130 trades and won 32.8% of them at a 2:1 target, which is a hair under the 33.3% it needs. That is −$37.90 per trade, −$459,730.54 in total, profit factor 0.878, t = −5.37. Expectancy is −0.018R, so the loss is friction, not a blown-up edge.
How much was the lookahead fill worth?
$2,750,727. Filling at the EMA price inside the bar whose close confirmed the pullback produced 50.1% wins, +$188.87 per trade and +$2,290,997 with a profit factor of 1.821. Moving the fill to the next bar's open, with the signals left untouched, turned that into −$37.90 per trade. Same rule, same 12,130 signals, one assumption apart.
Isn't the same-bar stop-and-target artifact the bigger problem?
Not on hourly bars with ATR stops. Only 166 of 12,130 trades, 1.4%, had both levels inside one bar. Scoring those as wins instead of losses moves the result from −$37.90 to −$24.85 per trade, worth $158,344 in total. It flatters the number, but the lookahead fill was more than seventeen times larger.
Does either side of the trade work on its own?
Neither, though the damage is lopsided. Longs took 6,598 trades at 34.2% wins for −$16.65 each, shorts 5,532 trades at 31.1% for −$63.25 each. The long side is the weaker loss, not a winner, and its t of −1.88 is not the signature of an edge hiding under the average.
Would a different target fix it?
No, and that sweep is the cleanest evidence in the study. At 1R the rule needs 50% winners and gets 47.3%; at 1.5R it needs 40% and gets 38.5%; at 2R it needs 33.3% and gets 32.8%; at 3R it needs 25% and gets 24.5%. Every configuration lands just under its own break-even line, which is what a fair coin minus $29.50 a round trip looks like.