Comparing Two Systems With Different Win Rates Fairly

Two rules, two win rates, and the obvious conclusion is that the higher one is better. The conclusion holds only if everything else about how the two numbers were produced was the same, and it almost never is. Most comparisons between trading systems are comparisons between the circumstances in which they were measured, wearing the clothes of a comparison between the systems.

Same Period, Same Instrument, Same Everything Else

Close-up of stock market trading screen displaying financial growth and charts.

The first requirement is the one most often skipped. A rule evaluated over a trending stretch and a rule evaluated over a choppy one have been given different examinations. Breakout approaches are particularly sensitive to this, since a period where ranges resolved into sustained moves flatters every variant of the idea and a period of repeated false breaks punishes them all.

Instrument matters equally. Different markets produce different opening range behaviour, different typical movement relative to that range, and different costs per trade. A rule measured on one and a rule measured on another have not been placed under the same conditions in any meaningful sense.

The honest version is to run both rules over the identical set of sessions and compare what each did on the same days. That removes the question of conditions entirely, and it usually shrinks the apparent difference considerably.

Hold the Exits Still, or Compare Them Deliberately

A hand points to colorful business charts and graphs on a paper sheet on a wooden desk.

If one rule uses a nearer target than the other, the win rate difference is partly an artefact of that choice rather than a difference in the entries. To compare the entry logic, both should be run to the same exits. To compare the exits, both should use the same entry. Changing both and reading one number is a comparison that cannot be interpreted.

This is worth being strict about because entry rules are what people care about and exit rules are what move the count. A comparison that leaves the exits free is very likely measuring the exits while everyone involved believes it is measuring the entries.

Frequency Changes What the Number Means

A rule that fires on most sessions and a rule that fires rarely can share a win rate and be entirely different propositions. The selective rule is doing more filtering, so its trades should be better, and if its win rate merely matches the permissive one, the filtering has bought nothing.

Frequency also determines how much confidence the number deserves. A count taken over a small number of trades will move a lot on the next few outcomes, and the difference between two such counts is likely to be noise. Comparing a figure drawn from a long record against one drawn from a short one gives an unearned appearance of equality between them.

Costs enter here too. A rule that trades often pays the spread and the commission often, and those costs fall on every trade including the winners. Two rules with similar gross behaviour can separate substantially once frequency is accounted for, and the win rate says nothing about it either way.

The Comparison That Actually Answers the Question

What people are usually trying to establish is which rule they would rather have traded. That question is answered by the whole distribution of outcomes rather than by the count, and it is answered best by looking at the record in a few pieces.

The size of the typical win against the size of the typical loss is the pairing that gives the count its meaning. The worst run of consecutive losses matters because it determines whether the rule is survivable. The shape of the equity curve over the period, in particular whether the gains arrived steadily or came from a handful of outsized days, decides whether the record describes something repeatable or something lucky.

A rule whose entire performance came from a small number of exceptional sessions is a rule you know very little about, whatever its win rate. The same is true in reverse: a rule whose losses are all one size and whose wins vary widely is behaving as designed, and the count is the least interesting fact about it.

When the Comparison Cannot Be Made

Sometimes the fair comparison is unavailable. The records were produced at different times, on different markets, by different people, and there is no way to reconstruct them onto common ground. The correct conclusion then is that the two cannot be ranked, which is unsatisfying and considerably more accurate than picking the larger number.

The habit worth building is to treat a comparison as a claim requiring evidence rather than as an observation. Asking what period, what instrument, what exits and how many trades takes a moment, and it will usually establish that the two figures were never comparable, which is itself the most useful thing you will learn from the exercise.