QFI Labs.
Reading the literature 5 minute read

Two thirds of published anomalies do not replicate

Hou, Xue and Zhang rebuilt 452 published anomalies under a single consistent method. Around 65 per cent of them could not clear the same significance bar the original papers had cleared.

The finance literature has produced a very large number of published return predictors. Harvey, Liu and Zhu counted 316 of them by 2012, and the rate has not slowed since. The set is large enough that it has a nickname, the factor zoo, and large enough that somebody had to go and check it.

How many published anomalies replicate A single bar split in two. Of 452 published anomalies re-tested under one consistent method, 294, or 65 per cent, failed to replicate. 158 held up. 294 failed to replicate 65% of the 452 tested 158 held up 452 published anomalies, re-tested under one consistent method Failure here means the effect could not clear a t statistic of 1.96 once every anomaly was measured the same way.
Each of the 452 anomalies was rebuilt from its original description and then measured the same way as every other: NYSE breakpoints for portfolio sorts, and value weighted rather than equal weighted returns. Source: K. Hou, C. Xue and L. Zhang, Replicating Anomalies, Review of Financial Studies 33(5), 2020.

The method is doing the work

This is not a case of fraud, or even of carelessness. Most of the failures come from two choices that look technical and are not.

Equal weighting gives a tiny company the same say as a large one. Since most of the stock market by count is small and most of it by value is not, an equal weighted portfolio can be driven almost entirely by names that nobody could trade in size. Value weighting removes that. So do NYSE breakpoints, which stop the smallest decile from filling up with microcaps.

Apply both, consistently, to every anomaly in the set, and two thirds of them stop being significant.

What we take from it

Before asking whether an effect is real, ask what would have to be true about the portfolio for it to be tradeable. If the answer involves buying a thousand illiquid names in equal amounts, the effect may be real and still be unavailable.

The bar itself is disputed

A t statistic of 1.96 is the conventional threshold, and it corresponds to roughly a one in twenty chance of a false positive on a single test. But the literature did not run a single test. It ran hundreds, published the ones that worked, and left the rest in a drawer.

Harvey, Liu and Zhu argue that once you account for how many factors were tried, the honest hurdle is closer to a t statistic of 3.0. Almost nothing published before 2012 was held to that standard.

How we use it

We treat a published anomaly as a hypothesis with a known prior, and the prior is not good. Anything we take seriously has to survive being measured the boring way: value weighted, with liquid names, with costs, over periods the author never saw.

Where the false positives come from