Blog
Thirty Trades Is Not a Track Record
After thirty trades and a very good month, I sized up. The arithmetic I ignored would have told me the data could not distinguish my edge from a coin flip.
I had a very good month once, early on. Thirty trades, a win rate somewhere in the sixties, and an average result comfortably above +0.4R per trade. I remember doing the multiplication to see what that would compound to over a year. Then I roughly doubled my position size.
The next two months were losing months. Not disasters, but losses, at the new size, which made them feel like disasters. I concluded that the strategy had “stopped working”, which is what traders say when they do not want to say “I never knew whether it worked”.
I write the statistics guides on this site now, so this post is partly an apology to my younger self for not doing the arithmetic, and partly the arithmetic.
What thirty trades can tell you
Expectancy is a sample mean. It has a standard error, like any sample mean, and the standard error shrinks with the square root of the number of trades. That last part is the whole problem.
Suppose the true edge of a strategy is +0.2R per trade, which would be a respectable retail edge. Suppose the R multiples have a standard deviation around 1.5R, which is normal when losers cluster near −1R and winners run to +3R or +4R. After thirty trades, the standard error is 1.5 divided by the square root of thirty, which is about 0.27R.
That means a thirty-trade sample of a strategy with a real edge will routinely show an expectancy anywhere from roughly −0.3R to +0.7R. My +0.4R month was squarely inside the range you would expect from luck operating on a modest edge. It was also squarely inside the range you would expect from luck operating on no edge at all. The data could not tell those apart, and I had not asked it to.
What I had done, in statistical terms
I had taken a noisy estimate, treated it as a precise one, and then increased the stakes on the strength of that precision. Sizing up after a good month is the most natural thing in the world and it is, mechanically, the same error as sizing down after a bad one. Both use a sample too small to justify any change, and both change the thing that most affects the outcome.
The bad months that followed were not the strategy failing. They were the strategy producing the other half of its distribution, which I had not seen yet because I had only watched it for thirty trades. Had I stayed at the original size, I would have had a flat quarter and a longer sample. Instead I had a losing quarter and a conclusion that was as unsupported as the one I had drawn after the good month.
The calculation I do now
Before I change anything about a strategy, size included, I compute three numbers from the journal.
The expectancy: the mean of the R multiples.
The standard error: the standard deviation of the R multiples divided by the square root of the trade count.
The interval: expectancy plus and minus about two standard errors.
Then I look at the lower bound. If the lower bound is below zero, the data has not shown me an edge, whatever the mean says, and I do not size up. If the lower bound is above zero, I have something, and even then I size to the lower bound rather than to the mean, because the mean is the optimistic reading and sizing to the optimistic reading is how the last version of me lost a quarter.
The expectancy and position sizing guide has the worked example, including why the Kelly formula makes this mistake worse rather than better if you feed it a small sample.
The uncomfortable implications
Two things follow that most traders do not want to hear.
First, most people’s trading history is too short to say anything. A hundred trades is the point at which the intervals start to be useful. Two hundred is better. Many retail traders change strategy every few months, which means they never reach a sample that could tell them whether any of the strategies worked.
Second, a strategy that “stopped working” after a bad month almost certainly did not stop working. It either never worked and you are seeing the truth, or it works and you are seeing variance. Thirty trades cannot tell you which. The only thing that can is more trades at a size you can survive, which is an argument for small size that has nothing to do with caution and everything to do with information.
What I would tell someone starting
Keep the size small for longer than feels necessary, and think of the small size as the price of a sample rather than as timidity. You are paying to find out whether you have an edge, and the payment is time at a size where variance cannot hurt you.
Do not react to any month. Do not react to any ten trades. Work out the standard error, look at the lower bound, and let that number make the sizing decisions. It will be slower than your instincts and considerably better at arithmetic.
And read backtesting without fooling yourself alongside this, because the same square root that makes live samples noisy makes backtests optimistic, and for the same reason.