> ## Content Index
> Fetch the complete content index at: https://aligrithm.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# 10.18 Moving Average Distance: The Technical Indicator That Passed the Cross-Section
- URL: https://aligrithm.com/moving-average-distance-the-technical-indicator-that-passed-the-cross-section/
- Published: 2026-09-23T04:54:03.000Z
- Updated: 2026-09-23T04:54:02.000Z
- Description: US equity cross-section: the 21-day over 200-day moving average ratio earns 9.05% annual alpha and kills momentum's. All of it is the long leg, and the short leg pays nothing.
- Author: ali askar
- Tags: 10. Cross-Sectional & Factor Investing

Divide a stock's 21-day moving average by its 200-day moving average. That is the whole signal, and it is about as retail as a signal gets. Avramov, Kaplanski and Subrahmanyam ran it across 13,828 US firms and 1,353,679 monthly returns from July 1977 through December 2018, and the value-weighted hedge portfolio produced an annual alpha of 9.05% against the five Fama-French factors plus momentum plus a crisis dummy, with a t-statistic of 3.02\. Momentum does not survive the encounter: the standard winner-minus-loser spread earns a 17.58% three-factor alpha on its own, and 2.25% with a t of 0.96 once you add the moving average distance factor to the right-hand side. One dollar placed in the top decile in July 1977 finishes at 959.69 dollars against 92.47 for the CRSP value-weighted market.

Now the part that decides whether you can trade it. The entire 9% sits on the long leg. The short leg's alpha is minus 0.05% a year with a t-statistic of minus 0.02, which is a rounding error wearing a minus sign.

## What this actually is

You own a trend filter. Someone told you that when the 50-day crosses the 200-day you buy, and when it crosses back you sell, and every clean test of that crossover on an index comes back flat or worse. The authors check this directly: a dummy for whether the golden cross or death cross fired this month carries a t-statistic of 1.17 in the 21-day version, which is nothing. The binary event is dead. The question is whether anything in the moving average pair is alive.

What they do about it: stop treating the crossover as an event and start treating the gap between the two averages as a number, then rank every stock in the market by that number each month and buy the widest gaps while selling the narrowest.

The mechanism they propose is anchoring. Move to a new city and your sense of what a coffee costs stays pinned to your old town for months. A price that is twenty cents off you accept on the spot. A price that is triple you dismiss as a tourist trap and only update grudgingly, over many visits. The distance between your running one-month impression and your running year-long impression measures how much updating you still owe. Their claim is that stock prices owe the same updates, and that the gap between the short and long averages is a readable meter of the debt.

At the desk this changes two things. You rank the cross-section on the ratio instead of trading crossovers, and you hold for three to twelve months instead of reacting to the cross. Believing it when it is wrong costs you specifically: the ratio correlates 0.58 with momentum and 0.64 with the 52-week high, so a book you think is a new anomaly may be a momentum book with extra turnover and extra fees.

## The ratio, and the screen that is supposed to sharpen it

The construction runs in two steps, and the second one does less work than it appears to.

$$ MA\_k(t) = \\frac{1}{k}\\sum\_{j=0}^{k-1} P\_{t-j}, \\qquad MRAT\_t = \\frac{MA\_{21}(t)}{MA\_{200}(t)} $$

Read it as: MA sub k at time t is the arithmetic average of the last k daily closing prices in dollars, split- and dividend-adjusted, and MRAT is the 21-day average divided by the 200-day average, a pure ratio with no units. A value above one says the recent month has traded above the year-long baseline.

Worked example. A stock whose last 21 closes average 52.00 dollars and whose last 200 closes average 45.00 dollars has an MRAT of 52.00 divided by 45.00, which is 1.1556\. A stock averaging 33.00 dollars over the month against 40.00 over the year has an MRAT of 0.8250\. Across the full sample the mean MRAT is 1.047 with a standard deviation of 0.204, so 1.1556 is a bit more than half a standard deviation above average and 0.8250 is about one standard deviation below.

The second step converts the ratio into a plus-one, minus-one, zero variable.

$$ MAD\_{i,t} = \\begin{cases} +1 & \\text{if } MRAT\_{i,t} \\in D\_{10} \\text{ and } MRAT\_{i,t} > 1 + \\sigma\_t \\\\\[4pt\] -1 & \\text{if } MRAT\_{i,t} \\in D\_{1} \\text{ and } MRAT\_{i,t} < 1 - \\sigma\_t \\\\\[4pt\] \\ \\ \\ 0 & \\text{otherwise} \\end{cases} \\qquad \\sigma\_t = \\operatorname{sd}\_i\\!\\left(MRAT\_{i,t}\\right) $$

Read it as: a stock scores plus one if it lands in the top MRAT decile that month and its ratio also clears one plus sigma, minus one if it lands in the bottom decile and falls below one minus sigma, and zero otherwise. Sigma is not a time-series volatility. It is the standard deviation of MRAT computed across all stocks in that single month, so the bar moves with how dispersed the market is.

Worked example with the sample-wide sigma of 0.204\. The long bar sits at 1.204 and the short bar at 0.796\. The stock at 1.1556 fails the long bar even if it makes the top decile, so it scores zero. A stock at 58.00 over 45.00, or 1.2889, clears it and scores plus one. The stock at 0.8250 misses the short bar of 0.796 and also scores zero.

That sounds like a tight double filter. Fit a lognormal to the Table 1 moments and it mostly is not.

![Density of the MRAT ratio fitted to the paper's Table 1 moments, showing the bottom decile cutoff at 0.803 nearly coinciding with the one-minus-sigma line at 0.796, and the top decile cutoff at 1.316 sitting above the one-plus-sigma line at 1.204](https://storage.ghost.io/c/27/cb/27cb0fc8-2c77-4434-af9e-d5d32a916994/content/images/2026/09/article_600-mrat_double_screen.png)

The decile cutoffs land at 0.803 and 1.316\. The sigma bars land at 0.796 and 1.204\. On the short side the two cutoffs are almost the same number, and on the long side the decile boundary is the binding one. Using the pooled dispersion, the sigma screen removes close to nothing, and "MAD equals plus one" is close to "top MRAT decile." The screen only bites in months when cross-sectional dispersion runs well above its sample average, which is exactly when the top decile is already extreme. The authors note that going to two sigma strengthens the results "at the cost of having several months with relatively small number of investable stocks," which is a polite way of saying the screen and the portfolio count trade off directly. Treat the sigma condition as a mild trim, not as the source of the edge.

## Surviving twenty-odd controls

The first test is whether the ratio still predicts once you throw the standard anomaly zoo at it. They use Fama-MacBeth, which means one cross-sectional regression per month and then a t-test on the time series of slopes.

$$ R\_{i,t+1} = a\_t + b\_t\\,MAD\_{i,t} + \\sum\_{k} c\_{k,t}\\,X\_{k,i,t} + e\_{i,t+1}, \\qquad \\hat{b} = \\frac{1}{T}\\sum\_{t=1}^{T} b\_t, \\qquad t(\\hat{b}) = \\frac{\\hat{b}}{\\operatorname{se}\_{NW}(b\_t)} $$

Read it as: for each of the 498 months, regress next month's return on the signal and on every control X, keep the slope b on the signal, then average those 498 slopes and divide by their Newey-West standard error. The controls here run past twenty variables, including size, book-to-market, idiosyncratic volatility, turnover, Amihud illiquidity, standardized unexpected earnings, net stock issues, asset growth, gross profitability, accruals, return on assets and equity, net operating assets, the Ohlson distress score, a dummy for price above the 200-day average, MACD, and four separate past-return windows.

The results. For one-month-ahead returns the continuous ratio carries a t of 4.92 and the plus-one-minus-one version carries a t of 6.65\. Over months two through six the slopes hold at t equals 3.60 and 6.12\. Over months seven through twelve both collapse, to t equals 1.51 and 0.66\. In the same regressions momentum turns negative and insignificant, minus 0.00 with a t of minus 0.03, and the 52-week high does the same at minus 0.76 with a t of minus 1.67\. The trend factor of Han, Zhou and Zhu is the one control that stays strongly positive alongside.

One arithmetic problem, and I am leaving it as the source printed it. Table 2's note says the slopes are multiplied by ten thousand. Take that literally and the MAD slope of 0.51 implies a top-minus-bottom spread of 2 times 0.51 times ten to the minus four, which is 0.0102% a month. Table 3 reports the actual raw spread at 1.58% a month. Reading the same slopes as percent per month gives 1.02% a month, which sits below the raw 1.58% by about the amount you would expect twenty controls to absorb. The scaling note is off by two orders of magnitude. The coefficients themselves are internally consistent, so nothing in the paper's conclusions moves, but do not quote the slope magnitudes without fixing the units yourself.

Two more things in this table deserve flagging. Harvey, Liu and Zhu argue that with thousands of candidate factors in circulation, a new one should clear a t-statistic of 3.0 rather than 2.0\. Most of the MAD t-statistics do. The 2001 through 2018 subsample does not: it comes in at 2.61 for MAD and 3.06 for the continuous ratio. So the claim that the effect persists into the modern era rests on a coefficient that fails the bar the same paper cites two pages later. And in negative market states, defined by a negative trailing two-year market return, the effect vanishes entirely, with t-statistics of 1.82 and 0.40\. The authors blame the small sample, and they are right that only 10.6% of sample months qualify, but the honest reading is that we have close to no evidence about how this signal behaves through a sustained bear market.

## The long leg is the strategy

Sort into value-weighted deciles and the raw spread is 1.58% for the next month and 7.27% cumulative over months two through six, both significant at 1%. Risk-adjust it and the picture splits.

$$ R\_{p,t} - R\_{f,t} = \\alpha\_p + \\beta\_1 MKT\_t + \\beta\_2 SMB\_t + \\beta\_3 HML\_t + \\beta\_4 RMW\_t + \\beta\_5 CMA\_t + \\beta\_6 UMD\_t + \\delta\\,Crisis\_t + \\varepsilon\_{p,t} $$

Read it as: regress the portfolio's monthly excess return, in percent, on the five Fama-French factors plus the momentum factor plus a dummy for 2008 and 2009, and call the intercept alpha. Alpha is the part of the return that the factor exposures do not explain, in percent per month; the paper annualizes by multiplying by twelve without compounding.

Worked example. The one-month top-minus-bottom portfolio has a monthly alpha of 0.7542%, and 12 times 0.7542 gives the 9.05% printed in Table 4\. That annualization convention matters when you compare against compounded numbers, because it overstates by whatever the compounding would add.

![Grouped bar chart of annual alphas by holding period for the hedge portfolio, the top decile alone, and the bottom decile alone, showing the long leg carrying essentially all of the alpha at every horizon](https://storage.ghost.io/c/27/cb/27cb0fc8-2c77-4434-af9e-d5d32a916994/content/images/2026/09/article_600-mad_alpha_by_leg.png)

At a one-month holding period the hedge earns 9.05%, the top decile alone earns 9.00%, and the bottom decile alone earns minus 0.05%. Out to twelve months the long leg actually grows stronger, at 8.26% with a t of 7.10, while the hedge decays to 6.67% because the bottom decile turns positive. Stambaugh, Yu and Yuan established that most anomalies live on the short leg, where arbitrage is expensive. This one inverts that, which is good news for anyone who cannot borrow shares and bad news for the story that the effect is mispricing that short sellers cannot reach.

The wealth curves make the same point in a way that also traps the unwary.

![Value of one dollar invested monthly from July 1977 to December 2018 in the MAD top decile, the CRSP value-weighted market, and the MAD bottom decile, on a log scale with NBER recessions shaded](https://storage.ghost.io/c/27/cb/27cb0fc8-2c77-4434-af9e-d5d32a916994/content/images/2026/09/article_600-mad_wealth_curves_1977_2018.png)

The bottom decile ends at 0.35 dollars, a loss of 65% over 41.42 years, which compounds to minus 2.50% a year. The top decile's 959.69 dollars compounds to 18.03% a year and the market's 92.47 dollars to 11.55%. So the short leg looks like a machine for losing money and yet its risk-adjusted alpha is zero. Both facts are true because the bottom decile is loaded with small, high-beta, high-volatility names whose losses the factor model already prices. That distinction is the whole point of the old article ["What a Factor Actually Is: α + βλ, and Why the Market Is Factor #1"](https://aligrithm.com/what-a-factor-actually-is-a-bl-and-why-the-market-is-factor-1/): a portfolio that loses money is not delivering negative alpha, it is delivering negative beta-times-lambda, and shorting it earns you the factor exposure rather than a premium.

## What it costs to actually run

Turnover is the reason most anomalies die on contact with a broker. The MAD top decile turns over 34.9% a month on average, the bottom decile 38.5%, which implies an average holding time of 1 divided by 0.349, or 2.87 months. The authors compute the transaction cost that would zero the alpha rather than assuming one.

$$ \\left\[R\_{l,t} - TO\_{l,t}\\times TC\\right\] - \\left\[R\_{s,t} + TO\_{s,t}\\times TC\\right\] \\ \\longrightarrow\\ \\text{regress on factors, solve } \\alpha(TC)=0 $$

Read it as: charge the long leg TC on every dollar it turns over and charge the short leg the same, both in percent per month, then find the TC that drives the regression intercept to zero. Everything in the bracket is a monthly percentage return.

Worked example. The three-factor-plus-momentum alpha at a one-month holding period is 8.35% a year, or 0.6958% a month. Combined turnover is 0.349 plus 0.385, or 0.734\. Dividing 0.6958 by 0.734 gives 0.948%, or 94.8 basis points, against the 94 basis points the paper reports. The five-factor specification gives 0.7542 divided by 0.734, or 102.7 basis points, against their 101\. The published long-short numbers reproduce.

The long-only row does not. The paper reports a break-even of 142 basis points for the top decile alone, but 9.00% a year is 0.75% a month, and 0.75 divided by the reported 34.9% turnover is 214.9 basis points. Charging a round trip instead gives 107.4\. Neither is 142\. The long leg's own turnover under that specification is never printed, so the figure cannot be checked. I am leaving it as published and flagging it as unverifiable rather than substituting my own number.

![Bar chart of break-even transaction costs by holding period against an institutional cost ceiling of 40 basis points and a retail one-way cost of about 208 basis points](https://storage.ghost.io/c/27/cb/27cb0fc8-2c77-4434-af9e-d5d32a916994/content/images/2026/09/article_600-breakeven_vs_real_costs.png)

Against Novy-Marx and Velikov's estimate of 10 to 40 basis points a month for trading momentum and post-earnings drift, every horizon clears comfortably. Against retail costs it is a different verdict. The value-weighted quoted spread averages 321 basis points for the top decile and 510 for the bottom, so a one-way cost near 208 basis points kills the one-month version outright and leaves the three-month version barely alive. The twelve-month version, with a break-even of 1,102 basis points, works for anyone. If you are trading your own account, the horizon is not a tuning parameter, it is the thing that decides whether the strategy exists. The old article ["Percentile-Rank Momentum With Hysteresis: Low-Churn Signals"](https://aligrithm.com/percentile-rank-momentum-with-hysteresis-low-churn-signals/) attacks the same problem from the signal side, by refusing to rerank on small changes, and that machinery transfers directly here.

## Anchoring, or a jump detector with a good story

The authors test their mechanism by splitting on suddenness. For each top-decile stock they take the largest positive monthly change in MRAT over the previous three quarters, and set SuddenUp to one if that change exceeds the decile's monthly median. SuddenDown is the mirror.

The results support anchoring in sign and undercut MAD in magnitude. SuddenUp carries a coefficient of 1.04 with a t of 6.66, SuddenDown minus 0.54 with a t of minus 3.01, and both point the way the theory predicts: the wider the jump away from the anchor, the more underreaction is left to harvest. Splitting the hedge portfolio, the sudden half earns 23.97% a year with an alpha of 12.57% at a t of 3.51, while the gradual half earns 14.30% with an alpha of 5.62% at a t of 2.60.

Here is what I take from that table and the authors do not spell out. MAD's own coefficient falls from 0.51 to 0.18 once the two suddenness interactions enter. More than half the effect is not "the averages are far apart" but "the averages got far apart in a hurry." That is a different signal with a different name. It is closer to a gap-and-drift detector, and it lands in the same territory as post-earnings announcement drift, though standardized unexpected earnings is among the controls and keeps its own positive slope at a t of 4.18, so the two are not the same thing. The overlap still deserves more than a footnote, and it means anyone implementing this should compute the suddenness split first and decide whether they want the whole signal or just the fast half.

For context on how large 9% is for a chart-derived quantity, the old article ["Price-Path Convexity: A New Cross-Sectional Anomaly (−45bp per σ)"](https://aligrithm.com/price-path-convexity-a-new-cross-sectional-anomaly-45bp-per-s-2/) covers a signal built from the shape of the price path rather than the level of two averages, priced at 45 basis points per standard deviation. MAD is a bigger number, computed from cruder inputs, on a longer sample. That asymmetry is itself a reason to keep checking it.

## The verdict

The result is real inside the sample and the authors earned it with a harder control set than most anomaly papers bother with. Three things keep me from sizing it like a discovered risk premium. The sample ends in December 2018, so the strategy has no out-of-sample record at all in the period since, which includes a pandemic crash, a zero-rate melt-up and a rate shock. The effect is untested in bear markets because 89% of sample months follow a positive two-year market return. And more than half the coefficient belongs to the suddenness interaction rather than to the distance itself. Run it long-only at a six- to twelve-month horizon, size it as a momentum variant rather than as an independent factor, and reserve judgment until someone posts the 2019 through 2025 numbers.

![](https://storage.ghost.io/c/27/cb/27cb0fc8-2c77-4434-af9e-d5d32a916994/content/images/2026/09/article_600-visual-0.png)

## KEY POINTS

- The distance between the 21-day and 200-day moving averages, not their crossing, is what predicts. The crossover dummy carries a t of 1.17 and dies; the ratio carries a t of 4.92 in a Fama-MacBeth regression against more than twenty controls, and the discretized version carries 6.65.
- The hedge portfolio's annual alpha is 9.05% against five Fama-French factors plus momentum plus a crisis dummy, and momentum's own alpha drops from 17.58% to an insignificant 2.25% once MAD enters. One dollar compounds to 959.69 in the top decile against 92.47 for the market over 41.42 years.
- All of the alpha is on the long side. Top decile 9.00%, bottom decile minus 0.05% with a t of minus 0.02\. The bottom decile loses 2.50% a year compounded, but that loss is factor exposure, not negative alpha, so shorting it pays nothing.
- Break-even costs reproduce exactly from the published turnover: 0.6958% monthly alpha divided by 73.4% combined turnover gives 94.8 basis points against the 94 printed. Institutions clear that at every horizon; retail, facing 321 and 510 basis point quoted spreads, only clears it past three months.
- Two numbers do not reconcile. Table 2's "multiplied by ten thousand" scaling note is off by two orders of magnitude against the portfolio spreads, and the 142 basis point long-only break-even does not follow from the published 34.9% turnover. Both are left as the source printed them.
- More than half the effect is a suddenness effect. MAD's coefficient falls from 0.51 to 0.18 once SuddenUp and SuddenDown enter, and the sudden half of the portfolio earns a 12.57% alpha against 5.62% for the gradual half.
- The sample stops at December 2018, the 2001 to 2018 subsample fails the t equals 3.0 bar at 2.61, and 89% of months follow a positive two-year market return. Nobody has shown this works in a bear market or after 2018.

## References

- [Moving Average Distance as a Predictor of Equity Returns - Doron Avramov, Guy Kaplanski and Avanidhar Subrahmanyam (SSRN 3111334)](https://papers.ssrn.com/sol3/papers.cfm?abstract%5Fid=3111334&ref=aligrithm.com)