# MIROSCOPE VARIANT BACKTESTS — ROUND 2 REPORT
Date: 2026-09-25. Dataset: `backtest/data/eurusd-h1.csv` (10,605 H1 bars,
2025-01-01 to 2026-09-22 UTC) — same as round 1. Engine: `variants/run_variants.py`
extended with round-2 classes (stdlib only, deterministic); round-1 classes
untouched and re-run to confirm exact reproduction. Accounting (all runs):
$10,000 start, 1% risk per setup, R-multiple accounting, no spread/commission/
slippage. Mechanization choices: `variants/PLAN-R2.md`.

## TOP LINE (plain language)
Round 2 produced the first positive results with triple-digit trade counts.
V3-L5 (RSI divergence, broad session): 120 trades, +0.502R expectancy, PF 2.76.
V3-L3 (looser 3-point divergence): 142 trades, +0.465R, PF 2.52. V2-EMA20:
107 trades, +0.378R, PF 2.24. The V3 divergence thesis survived a 9x increase
in sample size with expectancy intact — that is the single most important
finding of round 2. The 18-bar expiry was worth about +0.19R of round-1 V3's
number (V3-NX control), but the entry is still positive without it. The V1
strictness ladder revived V1 (n=30, +0.490R at RSI45) with a clean monotonic
curve. The V9 control flipped positive (+0.172R, n=52) under cleaner entry
mechanics, suggesting round-1's entry implementation, not the pattern, was the
problem. The V4 matrix shows a weak depth-0.25 ridge, nothing conclusive. The
V11 slope knob turned out to be non-binding (flat ladder). Nothing here is
out-of-sample, and nothing justifies paper trading yet.

## RESULTS TABLE (all round-2 runs + round-1 reference rows, ranked by expectancy then PF)

| Rank | Run | n | Win% | PF | Exp(R) | MaxDD% | Ret% |
|------|-----|---|------|----|--------|--------|------|
| — | V1a (r1) | 1 | 100.0 | inf | +2.250 | 0.0 | +2.25 |
| — | V3 (r1) | 13 | 84.6 | 5.13 | +0.635 | 1.0 | +8.25 |
| 1 | V3-L5 broad-session RSI div | 120 | 69.2 | 2.76 | +0.502 | 3.7 | +60.19 |
| 2 | V1-RSI45 loose reversal | 30 | 70.0 | 2.63 | +0.490 | 2.7 | +14.69 |
| 3 | V3-L3 3pt divergence | 142 | 66.2 | 2.52 | +0.465 | 3.9 | +66.02 |
| 4 | V3-NX no-expiry control | 13 | 76.9 | 2.92 | +0.442 | 2.0 | +5.75 |
| 5 | V1-RSI40 moderate reversal | 12 | 66.7 | 2.18 | +0.394 | 1.0 | +4.72 |
| 6 | V2-EMA20 reduced filter | 107 | 50.5 | 2.24 | +0.378 | 2.9 | +40.50 |
| — | V7 (r1) | 3 | 66.7 | 1.55 | +0.183 | 1.0 | +0.55 |
| 7 | V9c 50% pullback control | 52 | 42.3 | 1.39 | +0.172 | 4.0 | +8.96 |
| 8 | V4-T5D25 | 45 | 48.9 | 1.30 | +0.119 | 4.6 | +5.36 |
| 9 | V4-T2D25 | 31 | 48.4 | 1.27 | +0.117 | 4.3 | +3.63 |
| 10 | V11-A / V11-B / V11-C (identical) | 30 | 40.0 | 1.19 | +0.111 | 8.5 | +3.33 |
| 11 | V4-T3D25 | 39 | 48.7 | 1.26 | +0.105 | 3.9 | +4.09 |
| 12 | V4-T3D382 | 42 | 50.0 | 1.18 | +0.077 | 3.9 | +3.24 |
| 13 | V4-T2D382 | 35 | 48.6 | 1.17 | +0.076 | 4.3 | +2.66 |
| — | V11 (r1) | 23 | 39.1 | 1.12 | +0.068 | 7.2 | +1.55 |
| 14 | V4-T5D382 | 47 | 48.9 | 1.14 | +0.060 | 5.2 | +2.84 |
| 15 | V4-T3D15 | 27 | 48.1 | 1.08 | +0.039 | 3.1 | +1.05 |
| 16 | V4-T5D15 | 34 | 47.1 | 1.09 | +0.039 | 5.2 | +1.33 |
| 17 | V4-T2D15 | 24 | 45.8 | 1.06 | +0.031 | 3.5 | +0.73 |
| — | V4 (r1) | 24 | 45.8 | 0.94 | -0.029 | 4.2 | -0.69 |
| — | V5 (r1) | 16 | 37.5 | 0.89 | -0.068 | 4.9 | -1.09 |
| — | V2 (r1) | 24 | 50.0 | 0.79 | -0.088 | 4.6 | -2.10 |
| — | Phase 4 baseline (r1) | 100 | 31.0 | 0.78 | -0.139 | 27.6 | -13.82 |
| — | V9 (r1) | 10 | 30.0 | 0.40 | -0.396 | 4.7 | -3.96 |
| — | V1b (r1) | 1 | 0.0 | 0.00 | -1.000 | 1.0 | -1.00 |

(V1a/V1b excluded from ranking: a single trade each is not a test.)

## SENSITIVITY CURVES

### V1 strictness ladder (RSI-at-P1 threshold: 35 -> 40 -> 45)
| Threshold | n | Exp(R) | PF |
|---|---|--------|----|
| 35 (r1 V1a+V1b) | 2 | +0.625 | — |
| 40 | 12 | +0.394 | 2.18 |
| 45 | 30 | +0.490 | 2.63 |
Curve: monotonic. Loosening the binding constraint increased both trade count
(2 -> 12 -> 30) and expectancy held/rose. This is the opposite of an isolated
peak: the edge, if real, gets STRONGER as the filter loosens, which is what a
genuine pattern (rather than an overfit filter) should do. Caveat: the 35 point
used two narrow sessions (07-11, 13-17); 40/45 use 07-17, so the comparison is
not perfectly clean. n=30 is still small.

### V11 slope ladder (0.15 / 0.10 / 0.05 ATR)
All three rungs IDENTICAL: n=30, Exp(R)=+0.111, PF=1.19. The slope threshold
is non-binding: zero setups were rejected by the slope filter at any rung,
given the other V11 filters (MA origin + EMA alignment already select trending
bars). This is a flat curve from a dead knob, not evidence of a plateau. The
V11 result itself (+0.111, n=30) is marginally positive and consistent with
round-1 V11 (+0.068, n=23), but the ladder added no information.

### V4 retest matrix (expectancy; rows = timing bars, cols = depth ATR)
|       | D0.15 | D0.25 | D0.382 |
|-------|-------|-------|--------|
| T2    | +0.031| +0.117| +0.076 |
| T3    | +0.039| +0.105| +0.077 |
| T5    | +0.039| +0.119| +0.060 |
The D0.25 column is best in all three rows. That consistency is mildly
interesting (a weak ridge, not an isolated peak), but the deltas are small
(+0.03 to +0.12R), n is 24-47 per cell, and timing shows no pattern at all.
Honest call: suggestive, not established. Do not treat D0.25 as a finding
until it survives out-of-sample.

### V3 comparisons
- Divergence threshold 5 (L5) vs 3 (L3): +0.502 (n=120) vs +0.465 (n=142).
  Loosening added 22 trades at a cost of 0.037R. Stable plateau: the concept
  survives the looser threshold.
- Session 13-17 vs 07-17: not cleanly isolated (the broad-session runs also
  dropped the RSI(P1) filter per the spec's filter lists). The combined change
  took V3 from 13 to 120 trades with expectancy intact.
- ATR floor 0.60 (L5) vs 0.45 (ATR45): IDENTICAL results (n=120, +0.502R).
  The 0.45 floor binds on zero bars; the ATR knob is non-binding below 0.60.
- Expiry vs no-expiry: round-1 V3 (+0.635R, n=13, 18-bar expiry) vs V3-NX
  (+0.442R, n=13, no expiry). Same 13 setups fired. The expiry contributed
  roughly +0.19R — about 30% of round-1 V3's expectancy was the exit rule, not
  the entry. The entry is still positive (+0.442R) without it.

## PER-FAMILY VERDICTS
- **V3 family**: STRONGEST RESULT OF THE PROGRAM SO FAR. +0.50R on 120 trades
  (L5), +0.47R on 142 trades (L3). First variants with both positive expectancy
  and triple-digit samples. Divergence 5 vs 3 is a stable plateau. Expiry is a
  real but partial contributor. Deserves out-of-sample testing (more pairs,
  more years) before anyone acts on it.
- **V1 ladder**: revived. Monotonic strictness curve, +0.49R on 30 trades at
  RSI45. Keep on the watch list; needs more data.
- **V2-EMA20**: +0.378R on 107 trades, PF 2.24. Second independent
  triple-digit positive. The EMA20 + reduced-filter rework fixed what the
  60-MA channel got wrong. Watch list, out-of-sample next.
- **V9c control**: +0.172R on 52 trades vs round-1 V9's -0.396R on 10. The
  cleaner entry (stop at the 50% level) flipped the sign. Round-1's failure
  looks like an entry-mechanics problem, not a dead pattern. Weak positive;
  needs more data.
- **V4 matrix**: weak D0.25 ridge, nothing conclusive. The retest idea remains
  approximately breakeven across the matrix.
- **V11 ladder**: marginally positive (+0.111, n=30), slope knob dead.
  Unchanged verdict from round 1: not convincing.
- **Overfit check**: NO isolated single-rung peak anywhere. V1 monotonic,
  V11 flat (dead knob), V4 a gentle consistent ridge, V3 threshold stable.
  The pre-registered discipline held: we report curves, not a cherry-picked
  best cell.

## ASSUMPTIONS (mechanization choices — see PLAN-R2.md for the full list)
- "1 tick" = trigger + 0.01*ATR14 (round-1 used 0.05). Round-2 only.
- Pivots confirmed 2 bars late; indicators read at pivot bars only (no lookahead).
- V3-NX keeps round-1 V3's full filter set including RSI(P1)<=40/>=60; only the
  expiry was removed (per the spec's CHANGED table).
- V3-L5/L3/ATR45 follow the spec's filter lists literally: NO RSI(P1) filter.
- V11 ladder: no RSI filter (parent spec omits it).
- V4 retest touch band above P2 = 0.15*ATR (round-1 default, documented
  assumption); break-wait staleness 12 bars (round-1 default).
- V9c: impulse from closes, retracement on closes, entry stop at the 50% level,
  TP3 = signal-bar close +/- impulse.
- Pending (unfilled-setup) expiries: i3+8 bars for V3/V1/V2-EMA20, i3+6 for V11,
  retouch_bar+timing for V4 (spec-silent, round-1 defaults).
- One position per direction max. Session windows gate the entry-signal bar only.
- Round-1 classes re-run and reproduce round-1 results exactly.

## COMPARABILITY CAVEAT
Same data, same $10k/1% R-multiple accounting as round 1, but round-2 runs use
a 0.01*ATR tick buffer vs round-1's 0.05, and several round-2 specs dropped
round-1 filters (RSI(P1) in V3-L5/L3/ATR45, RSI band in V11, percentile filter
in V1). Cross-round comparisons are close but not exact. All round-2 runs are
in-sample on the same 20-month EURUSD window as round 1. Nothing here has seen
out-of-sample data.

## BOTTOM LINE FOR MICHAEL
Three things survived with real sample sizes: V3-L5 (+0.50R, 120 trades),
V3-L3 (+0.47R, 142 trades), V2-EMA20 (+0.38R, 107 trades). Two more are alive
but thin: V1-RSI45 (+0.49R, 30 trades), V9c (+0.17R, 52 trades). The sensitivity
curves show plateaus and monotonic improvement, not isolated peaks — that is
the honest shape of maybe-something rather than overfit noise. The next honest
step, if he wants it: out-of-sample data (more pairs, more years) on V3-L5/L3
and V2-EMA20. That is his call; it is not started. Nothing here justifies paper
trading.
