The method behind every verdict · updated 28 September 2026
We test trading strategies the way your account would live them. Every issue follows the same steps, in this order.
Rules first. We rebuild each strategy from its author's own rules and write down every assumption. Then we reproduce the author's claim on the author's own window. If it does not reproduce, we say so and ask the author.
Clean data, on the exchange's calendar. Daily tests use ETF prices with dividends reinvested. Intraday tests use Dukascopy's free 1-minute index CFD prices, bid side, priced like the matching futures contract. Every daily series runs on the New York Stock Exchange calendar, with sessions taken from QQQ's and SPY's own price history. In minute data we first remove placeholder minutes: flat candles where open, high, low and close are equal and volume is zero. We take no entries on days the NYSE was closed.
Costs, then more costs. Every trade pays commission and slippage. Then we double and triple them. An edge that dies at twice the cost was never there.
The account, not the trade. A great average trade can still make a poor account. We run every strategy as an account whose idle cash earns the 4-week Treasury bill rate, credited daily. We compare it with holding QQQ, the bar, and SPY, the floor, over the same days: total return, dividends included, from each ETF's first trading day where the test allows.
Years it never saw. Most strategies were built on one stretch of history. We split the data into the years a strategy was tuned on and the years it was not, and compare each half with the index in the same half.
Next-door settings. We nudge every parameter. A real edge still works one step either side, in both halves. A lucky setting works at one value only.
Count the tries. Test ten versions and one will look good by chance. We count every version and raise the bar to match: with ten tries, t = 2 becomes about 2.8.
Real fills. A backtest fills the instant a bar closes, and nobody can. We re-price entries from 0 to 300 seconds late. Our own automated orders fill about 47 seconds after the signal.
The Retail Gate. Could someone with a day job and an ordinary broker account run it as written? If not, we say what it would take.
Numbers from files, checked twice. A script regenerates every number we publish from stored result files. Nobody types a number into a table. Before anything goes out, a separate check re-derives every headline number. Before we change our test engine, a regression job must reproduce every number we have already published.
How we compute the figures
Yearly figures use elapsed calendar time: days divided by 365.25. Never a count of rows.
Sharpe ratio: the average daily return in excess of T-bill cash, divided by its sample standard deviation (n − 1), times the square root of sessions per elapsed year. On the NYSE calendar that is √252.
Precision: we store every result at full precision and round once, when we display it.
What the verdicts mean
Fails: no edge after costs.
Loses to the index: the edge is real, but the account trails simply holding QQQ or SPY.
Unproven: the data cannot yet separate a real edge from luck or tuning.
Replicates: the author's claim reproduces on our data.
Works as an add-on: it trails the index alone but helps beside it. We say how strong the evidence is.
Fairness and corrections. We name and link every author, lead with what replicated and offer a right of reply before publishing. When a published number changes, the post gets a dated "Updated" note at the top that says what changed. When a verdict changes, we say so plainly, in the post and in every verdict so far.
Disclaimer. Past Performance Lab discusses how trading strategies behaved on historical data. It is not a financial advisory service, and the author is not a licensed financial adviser. Nothing here is a recommendation to buy, sell or hold any security or to follow any strategy, and nothing is tailored to any reader's circumstances. When we test a strategy a reader sends us, we publish the result for everyone; we do not answer individual investment questions. We accept no payment from the authors or products we test. Backtested results are hypothetical, and readers should not rely on this publication in making an investment decision.
The same disclaimer is on our About page.

