Outcome Bias: A Profitable Trade Is Not Necessarily a Good Decision
You moved a stop, price recovered, and the trade ended in profit. Was it a good trade? The account gained, but luck may have produced that gain. Outcome bias means judging a decision's quality by its result, and it can become an expensive trading habit.
What is outcome bias?
Outcome bias evaluates a decision using what happened afterward, rather than the information available when it was made. The same decision is remembered as good when it works and bad when it does not.
In probabilistic situations, that scoring method is unreliable. Good decisions lose and bad decisions win. One outcome is only loosely connected to decision quality. Yet a profitable account day can feel like approval of everything done that day, including rule violations. That is the central problem.
This pairs with survivorship bias. Survivorship bias concerns who is missing from the sample; outcome bias concerns how included observations are judged. Both can introduce information unavailable at entry into a later evaluation.
A two-by-two scorecard separates decisions and results
Decision quality and outcome are separate axes, creating four categories.
① Good decision + Good result → Justified win
② Good decision + Bad result → An unavoidable loss; behavior worth continuing
③ Bad decision + Good result → A dangerous win; the most costly category
④ Bad decision + Bad result → An unsurprising loss
→ The original guide's following line groups ①② as “wins” and ③④ as “losses” when describing outcome-only scoring.
→ Its central point is that outcome bias teaches people to abandon ② and repeat ③.
Category ③ is costly because the reward arrives immediately. If moving a stop turns a would-be loss into a profit, the brain rewards moving stops. That behavior repeats until it produces a large loss, potentially exceeding all the losses previously avoided.
The expected value of moving a stop
A numerical example shows why category ③ can be dangerous. Assume a habit of canceling the stop and holding when price reaches it.
Planned stop: −1R, or 1% of the account
After cancellation: 80% recover; 20% continue falling
Recovery, 80%
Avoid −1R and finish at +0.5R → +0.5R
Continued decline, 20%
Exit well beyond the stop at −5R → −5R
Expected value = 0.8 × (+0.5) + 0.2 × (−5)
= +0.4 − 1.0 = −0.6R
→ Eight out of ten times feel like a good decision.
→ Across twenty repetitions, the expected total is −12R.
An 80% win-rate habit can erode the account. It is difficult to notice because the eight recoveries feel brief and pleasant, while the two painful losses may be dismissed as bad luck. In this example, the behavior has negative expectancy, so repeating it produces an expected loss.
A rule-following −1R stop belongs in category ②: a poor outcome from a good decision. Outcome-only scoring calls it failure and encourages postponing the next stop. This creates a system that punishes good behavior and rewards bad behavior.
One trade tells you very little
Results tempt us because they arrive immediately. But a single probabilistic outcome contains far less information than it appears to.
Five or fewer wins: stated as around 50%
Three or fewer wins: stated as about 10%
Ten consecutive losses: 0.45 to the tenth power ≈ 0.034%, or about one in 2,900
→ The guide describes nearly half of ten-trade windows as looking like losing periods even for this strategy.
→ Changing strategy then can mean abandoning a good decision process because of a bad short-term result.
Small-sample results contain substantial noise. As discussed in losing-streak probability, long losing sequences can occur even under a normal strategy. Scoring solely by short-term results can mistake noise for a signal and cause repeated strategy switching, potentially encountering each new strategy's bad period afresh.
The necessary sample varies by strategy, but the guide recommends postponing conclusions with fewer than thirty trades. While building the sample, inspect rule compliance, which is directly observable even with few trades.
Record decisions, not just results
A trading journal containing only PnL cannot diagnose outcome bias. Reviewing it later shows only which trades made or lost money. Record decision quality in a separate field.
① Entry rationale, using only information visible at entry
② Stop location and why it was chosen
③ Target and expected risk-reward ratio
④ Rule compliance: followed or violated, with the violated rule identified
⑤ Outcome: PnL and R-multiple
→ Score conduct using ④ and aggregate performance using ⑤ separately.
→ Specifically count trades where ④ is a violation but ⑤ is positive.
Cross-tabulating ④ and ⑤ creates the two-by-two scorecard from your own records. A key measure is the number of profitable rule violations. If it grows, habits may be deteriorating even while the account gains. Category ②, following rules but losing, is normal; changing rules simply to eliminate that category can lead toward ③.
Deleting rule-breaking trades is another common mistake. Removing a trade because it was an error leaves its PnL in the account but removes the evidence of how often rules were broken. Overtrading and revenge trading often disappear from journals this way.
The same problem when judging others
Outcome bias can be stronger when viewing other traders because results are visible while their decision processes are not.
Account A: +180% in one month
· 2% of equity risked per trade, sixty trades, maximum drawdown −14%
Account B: +180% in one month
· 40% of equity committed at once, three trades, maximum drawdown −55%
→ Returns look identical.
→ B's approach could lose the account on a fourth trade.
Return alone does not reveal decision quality. Relevant context includes position size, maximum drawdown, trade count and record length. Without those four, a performance screen cannot distinguish category ① from ③. The original guide argues that exceptionally large short-term returns warrant particular attention to the dangerous-win possibility.
Apply the same standard to yourself. Repeated simulations using the same rules, such as Monte Carlo simulation, can place current performance within a range of possible outcomes. If it lies at the very top, the guide recommends allowing for a substantial contribution from luck.
Common misunderstandings
“Results tell the whole story.” Over a sufficiently large sample, results are informative. The problem is that sufficient may mean dozens or hundreds of trades, while judgments are often made after one or two.
“Focusing on process means being satisfied with losses.” Process evaluation does not justify every loss. Distinguish category ②, following rules and losing, from category ④, breaking rules and losing. That is why there are four cells.
“Adapting flexibly is a skill.” Responding through predefined exception rules can be skillful. An exception invented on the spot is a rule violation and should be recorded as such even when profitable.
“A good backtest proves my decisions were good.” A backtest evaluates rules, not every decision you actually made that day. Deviations made in front of the screen are absent from the backtest.
Key points
② Bad decision plus good result is especially costly because immediate reward encourages repetition.
③ Postponing stops can have negative expectancy even with an 80% win rate.
④ Ten-trade results can be dominated by noise; inspect rule compliance.
⑤ Record compliance separately and count profitable violations.
Score the account by outcomes and habits by decisions. Combining them approves poor habits during good results and discards good habits during poor results. Habits are what persist.
NOONOO TRADING invites you to follow live trading in our free chat.
Start in the bot📈 OKX trading fee discount for new registrations
Register for the OKX Fee Discount →