When Sharpe Lies: Evaluating Strategies Under Non-Normal Return Distributions

Joshua Goldfein · Apr 28, 2026
Joshua Goldfein · Apr 28, 2026

Two strategies arrive at a review with the same Sharpe ratio. The first has spent its life grinding out small gains against shallow give-backs. The second has done the same thing, plus one stretch where the position moved far enough against it to put the account itself in question. Sharpe reports them as equivalent, and it is behaving correctly when it does. It was built to compress a return distribution into two moments and report their ratio, and equivalence in those two moments is all it ever claimed.

When returns are approximately normal, mean and standard deviation carry most of the information in the series, and that compression costs almost nothing. Fat tails and skew change the accounting. The moments Sharpe keeps are the ones that describe the middle of the distribution, and the region it discards is where the strategy either survives or does not.

What the denominator can see

Sharpe measures excess return per unit of total volatility, and total is the operative word. The denominator scores upside deviation and downside deviation on identical terms, so a large favorable outlier is charged at the same rate as an unfavorable one of the same size. For a symmetric, thin-tailed series that is defensible, since deviation in either direction really is the same kind of information about the process. For a skewed series it collapses two directions with different consequences into one number and then reports the average as risk.

Standard deviation carries a second property that is easy to forget under pressure: it is order-independent. Twenty losing days scattered across a year and the same twenty arriving consecutively produce identical variance and therefore identical Sharpe. They do not produce identical margin treatment, financing cost, redemption behavior, or operator judgment at the moment of decision. Sequence is invisible to a period statistic, and sequence is where most books actually break.

Heavy tails compound both effects through the sample itself. In a quiet window the large moves that would raise the denominator have often not arrived yet, so the observed volatility understates the dispersion the process can produce. A small denominator yields a high Sharpe, and a high Sharpe reads as permission to size up. The number can look strongest just before a tail that the sample has not seen.

Four denominators, four questions

Sortino, Calmar, and Omega are worth treating as different questions rather than better answers. Each one changes the denominator, and the denominator is a declaration about which loss the evaluation is willing to see.

  • Sharpe

    — excess return per unit of total volatility. How much return per unit of overall dispersion?

  • Sortino

    — excess return per unit of downside deviation. How much return per unit of the dispersion that actually hurts?

  • Calmar

    — return relative to maximum drawdown. How much return per unit of the worst peak-to-trough path the book has already walked through?

  • Omega

    — probability-weighted gains above a stated threshold against probability-weighted losses below it. Relative to a return level declared in advance, how does the whole distribution trade off?

The differences between them are structural rather than cosmetic. Sortino keeps a dispersion denominator and restricts it to one side of the distribution. Calmar abandons dispersion for a path statistic, which makes it the only one of the four that responds to sequence. Omega abandons a single summary moment and integrates both sides of a threshold, which lets information from the tails reach the result instead of being averaged out two moments earlier.

None of that establishes a ranking. A drawdown denominator is right when the loss being underwritten is a peak-to-trough path and useless when the sample is too short for maximum drawdown to mean anything. A threshold ratio is precise when the threshold is a real obligation and arbitrary when the threshold was chosen after the results were visible.

Matching the denominator to the loss

The operating rule is short. Choose the ratio whose denominator matches the loss you are underwriting.

Where the binding constraint is a peak-to-trough path — capital withdrawn at a stated level, leverage cut by a risk desk, a mandate pulled — the evaluation belongs in drawdown terms. Where the constraint is dispersion below a required rate, the evaluation belongs on one side of the distribution and anchored to that rate. Where the constraint is ruin rather than discomfort, none of the four ratios is sufficient alone. Sharpe and Sortino still summarize dispersion, Calmar still records a path already walked, Omega still integrates around a threshold, and ruin lives in the region those summaries leave behind.

A public-safe system would make that a design question asked before anything gets ranked: what loss ends this book, and can this metric see that loss? A metric that cannot see the terminal loss is measuring comfort.

Where evaluation stacks fail

Four failures recur, and all four are procedural rather than statistical.

  1. The first is metric shopping.

    With four ratios available, one of them usually flatters the strategy, and the discipline that prevents this is declaration order: name the evaluation set before the run and report every declared metric whether or not it agrees with the others.

  2. The second is sample-length theater.

    Maximum drawdown is monotone in sample length, so Calmar computed on a short window describes the window rather than the strategy. Downside deviation and Omega fail in the opposite direction, since on short samples the estimate is dominated by a handful of observations and unstable in precisely the region the evaluation cares about.

  3. The third is comparison across unlike shapes.

    Ranking two strategies by any of these ratios assumes their distributions are comparable objects. Placing a high-kurtosis series and a near-normal series on one axis produces a winner and destroys the information that would have made the winner meaningful.

  4. The fourth is threshold smuggling.

    Omega’s value depends entirely on the threshold, and a threshold that travels unpublished alongside the number makes the number unfalsifiable.

Public surface, private book

The public part of this work is the evaluation framework: what each denominator can see, which loss each is fit for, and the procedure that keeps the metric set honest. That vocabulary is standard, and it travels to any domain where a summary score stands in for a risk.

The private part is every book-level input and output — the strategies under evaluation, the parameters that generate them, the thresholds and windows actually used, and every resulting value. Those stay internal.

This note describes an evaluation design rather than a result. AlphaFlux publishes no Sharpe, Sortino, Calmar, or Omega figure here, claims no track record, and describes no shipped evaluator. Nothing here is investment advice.

The test worth keeping

Before any ratio is allowed to rank anything, ask which part of the return distribution the metric can see, and whether that is the part that ends the strategy.

Sharpe answers a real question about dispersion, and under non-normal returns it answers that question while staying silent on skew, kurtosis, and sequence. The correction lives in the structure of the evaluation rather than in a better single number: denominators chosen deliberately, declared before the run, and matched to a loss someone has agreed to carry.


Image
AUTHOR NAME

Joshua Goldfein

Joshua Goldfein is a digital strategist with 20+ years of experience leading global teams, launching high-impact digital products, and driving growth through innovation, systems thinking, and AI integration.