Trade Sample Size: How Many Trades Are Needed to Assess a Win Rate?
Twelve wins and eight losses in twenty trades gives a 60% win rate. It looks like evidence that a strategy works, but the original guide notes that such a result can still arise from a much lower underlying win rate. The sample needed for a judgment can be estimated mathematically.
A win rate has an error range
Twelve heads in twenty coin tosses does not establish a biased coin. Likewise, a 60% win rate over twenty trades is one sample, with sampling uncertainty.
For win rate p and trade count n, standard error = √(p × (1−p) ÷ n). The guide uses the approximate 95% interval observed win rate ± two standard errors.
n = 20: SE = √(0.6 × 0.4 ÷ 20) = 11.0 percentage points
→ Approximate interval 38%–82%
n = 100: SE = √(0.24 ÷ 100) = 4.9 points
→ Approximately 50%–70%
n = 400: SE = √(0.24 ÷ 400) = 2.4 points
→ Approximately 55%–65%
With twenty trades, the lower endpoint is 38%, leaving substantial uncertainty about whether the underlying method is profitable. At 100 trades, the lower endpoint only approaches 50% in this approximation.
Halving the standard error requires four times as many trades. Increasing from twenty to 100 trades multiplies the sample by five but only reduces the error from 11.0 to 4.9 percentage points. A modest addition to the sample adds less certainty than it may feel.
Fees raise the breakeven win rate
50% may feel like the natural baseline, but reward-to-risk and costs determine the actual threshold. Even with equal gross wins and losses, fees push it above 50%.
Gross target and stop each produce ±$10.
Round-trip fees 0.1% → 1,000 × 0.1% = $1.
Expectancy = p × 10 − (1−p) × 10 − 1 = 20p − 11.
Set to zero → p = 0.55.
Breakeven is 55%, not 50%.
At 53% over 300 trades:
Expected PnL = 300 × (20 × 0.53 − 11) = −$120.
A 53% win rate sounds winning but loses under these assumptions. Using 50% as the threshold can keep a trader committed to a losing method. Calculate the relevant breakeven first; see win rate and reward-to-risk and expectancy and R-multiples.
Smaller edges require much larger samples
Required sample size depends on the difference you want to distinguish from uncertainty. The guide's approximation solves for n where two standard errors are smaller than the difference of interest.
Use p(1−p) ≈ 0.25 for convenience.
A five-percentage-point difference, such as 55% versus 50%:
Required SE 2.5 points → n = 0.25 ÷ 0.025² = about 400 trades.
A three-point difference, such as 58% versus 55%:
Required SE 1.5 points → n = 0.25 ÷ 0.015² = about 1,100 trades.
A two-point difference:
Required SE one point → n = 0.25 ÷ 0.01² = about 2,500 trades.
A small edge can be difficult for an individual to validate through live trading. At three trades a day, the original guide estimates roughly four months for 400 trades and more than two years for 2,500. Market conditions can change in the meantime.
The guide therefore favors focusing validation on fewer, stronger candidate setups rather than increasing frequency merely to collect observations. It warns that higher frequency also increases fee expenditure and describes that as raising the cost hurdle.
Three habits that damage a sample
① Selecting only a good period. “The last thirty trades won 70%” is misleading if that interval was chosen after seeing the result. Favorable runs can appear by chance. Define the period before observing its outcomes.
② Testing many variants and keeping only a winner. Trying twenty combinations of stops, indicator periods or sessions can produce an apparently successful variant by chance even when none has an edge.
In the guide's calculation, the probability that at least one passes when all have no edge is:
1 − 0.95²⁰ = 1 − 0.358 = about 64%.
→ More than half the time, an apparent “discovery” appears.
More variants require stricter standards.
This occurs often in backtesting. Repeatedly changing conditions until the curve looks attractive can amount to memorizing historical data.
③ Counting trades without considering order. Identical win rates and payoff ratios can produce very different account paths when losses cluster. Monte Carlo simulation explores alternative sequences, while losing-streak probability explains how long sequences can occur normally.
What to assess while the sample is small
A small sample does not prevent every useful judgment. Change what you evaluate.
Inspect execution rather than win rate. Even twenty trades can show whether entries met the rules, stops were moved or sizing was followed. These concern the process rather than uncertain returns. A trading journal supports this review.
Inspect the loss distribution. The largest loss, the count of losses beyond the planned stop and the depth of losing sequences can flag risk even in small samples. The original guide treats any loss larger than planned as a rule-compliance concern rather than simply a sample-size problem.
Write the decision rule in advance. For example, “At 300 trades, retain only if net PnL after fees is positive.” Specifying the number and evaluation point first prevents selecting a favorable window afterward.
Four practical checks
① Calculate breakeven for your own conditions. Use the threshold incorporating fees and payoffs, rather than assuming 50%.
② Set the sample size before starting. The guide's rough calculation requires about 400 trades for five points and 1,100 for three points.
③ If rules change, start the sample count again. A changed stop distance defines a different strategy; do not merge earlier trades into the new rule's sample.
④ Check sample size especially carefully when results look good. Favorable results are often accepted more readily than poor ones, which makes them particularly important to scrutinize.
Three key points
① Win rates have uncertainty. The approximate interval around 60% from twenty trades is 38%–82%, too broad to distinguish many profitable and unprofitable possibilities.
② Fees raise the breakeven threshold above 50% in the equal-payoff example. A wrong baseline can make a losing method look profitable.
③ Sample requirements grow with the inverse square of the difference: about 400 trades for five points and 2,500 for two points in the approximation.
Caution
Win rates, fees, leverage and payoff sizes are hypothetical calculation examples, not measured strategy performance. The standard-error formula assumes independent outcomes and a constant win probability. Real win probabilities change with market conditions, so the calculated counts are closer to lower-bound planning estimates in the guide's interpretation. Leveraged trading can lose all principal. Decisions and their consequences remain your responsibility.
NOONOO TRADING invites you to follow live trading in our free chat.
Start in the bot📈 OKX trading fee discount for new registrations
Register for the OKX Fee Discount →