Every headline number on this site, the threshold it was judged against, and the bars it does not clear. Nothing on this page is typed by hand — it is read from the backtest artifact and the placebo study in the repository, so it cannot quietly go stale.
The obvious objection to any backtest is that the machinery itself manufactures the result. So we destroyed the signal and kept the machinery: within each rebalance date the scores were randomly reassigned between companies — preserving the distribution, the missing data and the correlation between themes, destroying only the link between a company and its own score — and the whole pipeline was run again, 100 times, on seeds 1000..1099.
Grey is where those 100 deliberately worthless runs landed. The dark line is the real result.
2 of these beat every single noise run outright. The rest did not, and the counts are on the labels above rather than rounded away: long-short spread, t-statistic was matched or beaten by 3 of 100, decile ladder ordering was matched or beaten by 1 of 100. That is why the long-only ranking is the claim this site leads with and the long-short spread is not.
PLACEBO_HAC.json,
which retains all 100 draws
(pre-registered as PREREG_ma19_ma13_recalibration.md).
The "noise bar" is the 95th percentile of those draws. A tie counts against us: a noise run
that merely equals the real result is counted as having matched it.The top-tenth ranked by our score, rebalanced quarterly, held for one quarter, equally weighted. Measured against three different comparisons rather than the single most favourable one.
| Benchmark | It returned | Top decile | Difference | t | Quarters ahead |
|---|---|---|---|---|---|
| equal-weight universe (incumbent) uninvestable — you cannot buy this every name in the panel, cost-free — uninvestable, the number every historical alpha figure in this project used |
18.1% | 25.3% | +7.17% | 4.38 | 71% |
| cap-weighted panel average
closest investable analogue buildable from the panel itself |
14.9% | 25.3% | +10.46% | 4.29 | 68% |
| SPY total return over the same windows
what the user's obvious alternative actually returned |
15.3% | 25.3% | +9.99% | 3.77 | 64% |
All figures annualised and gross of trading costs. The equal-weight universe is listed first because every historical figure in this project was measured against it — and it is the hardest of the three, because an equally weighted basket of every company in the panel, charged nothing to trade, beat the S&P over this window. The SPY row is the one that answers "versus just buying the index".
A top decile can look good by luck. The question is whether the whole ranking orders returns. Annualised return by score decile, best-ranked first.
It sorts, and imperfectly — deciles 4 through 7 are effectively tied and one is out of order. The ordering statistic is -0.891, where −1.0 would be a perfect ladder. What carries the result is the top decile and the bottom two; the middle is noise.
An annualised average is one number standing in for 69 quarters, and quoting it alone is how a strategy that is often behind gets described as though it were always ahead.
"Net of costs" asks you to believe our cost estimate. The breakeven does not: it is the cost at which the edge reaches exactly zero, so you can compare it to whatever you think you would really pay.
After costs the top decile's edge over the equal-weight universe falls from 8.12% to 6.07% a year.
Four standard thresholds. It clears one and fails three. All four are here because a page that showed only the passing one would be advertising, not evidence.
| Threshold | Result | Bar | |
|---|---|---|---|
| Long-short t vs. this project's own placebo floor The 95th percentile of 100 shuffled-signal runs. Beating it means the result is bigger than 95 out of 100 runs on a signal known to be worthless. |
2.62 | 2.284 | clears |
| Long-short t vs. the Harvey-Liu-Zhu multiple-testing hurdle A bar that rises with the number of tests you have run. We have run a lot. This is the headline's clearest failure and the artifact records both sides of the argument. CLEARS the bar measured against the project's own placebo and FAILS the bar derived from counting its own trials. Neither is 'the' answer. HLZ prices the best of N draws; the deployed composite is flat 1/7, never tuned, and cpcv.adopt is false on every run, so the logged trials are overwhelmingly REJECTED ALTERNATIVES to it rather than candidates it beat. |
2.62 | 3.29 | fails |
| Probability of backtest overfitting Fails. And the bar is close to useless here: on a signal shuffled into pure noise this statistic reads about 0.47 on average, so roughly half of all worthless signals 'pass' it. We report it failing rather than quietly dropping a measure that does not flatter us. |
0.733 | 0.5 max | fails |
| Deflated Sharpe ratio Fails the conventional 0.95 bar. It is a genuine deflated figure — it is charged every one of the tests below, not the eight a naive version would use — and it sits above all 100 shuffled runs. Both halves are true and we do not quote one without the other. |
0.786 | 0.95 | fails |
The reason a backtest is weak evidence is that you can keep trying until something works. The defence is to write down every attempt, including the ones that failed, so the successes can be discounted by how many shots were taken. Every test below was registered in writing — hypothesis, threshold and direction — and committed to version control before it was run.
None of these were discovered by an outside auditor. They were found, written down and published by the project itself, which is the only reason you are reading about them.
Four things you can verify without any special access, in rough order of how much effort they take.
Every figure above is read at request time from
BACKTEST_RESULTS.json and MA19_RECALIBRATION.json. Nothing is
typed into this page. The run is stamped
57bc3f2, dated
2026-08-14 — if that date is old, these numbers are old,
and this page will say so rather than quietly refresh the label.
The SPY row above says the S&P returned 15.3% a year over 2009-01-15 to 2026-01-28. That is a public number you can look up in thirty seconds. If our benchmark is wrong, everything measured against it is wrong, and it is the easiest thing on this page to falsify.
The usual killer is look-ahead: scoring a company in 2012 using a number that was not published until 2013. This panel is rebuilt as of each date from filings with their real filing dates, prices with no future adjustment, and institutional holdings lagged by their actual disclosure deadline — and companies that went bankrupt or were delisted stay in, so the record is not a survey of survivors. The How it works page explains the construction and the specific places it is still imperfect.
Each one was registered before it ran, with the threshold written down first, and the verdict recorded whichever way it came out. That record is what makes the 248 figure above meaningful rather than decorative — a count of attempts is only honest if the failures were logged at the time rather than reconstructed afterwards.
The most convincing test this project has run is an out-of-sample replication on international data: the same untuned model, mapped onto a different vendor's dataset, different countries, a different construction and a different period. It works there, and the United States is the weakest region tested. Those figures are derived from a dataset licensed for non-commercial research only, so they cannot be published on a product page. That is a licensing constraint, not a hedge — and a page claiming to show proof should say what it is leaving out and why.