The evidence

Every headline number on this site, the threshold it was judged against, and the bars it does not clear. Nothing on this page is typed by hand — it is read from the backtest artifact and the placebo study in the repository, so it cannot quietly go stale.

Backtest run 2026-08-14 Commit 57bc3f2 2,531 companies 69 quarterly rebalances 2009-01-15 to 2026-01-28
These are backtested results, not a track record. They describe what the ranking would have done on historical data that was reconstructed to contain only what was publicly known on each date. They are not money that was made, they are hypothetical, and a backtest is the weakest form of evidence there is. Everything below is arranged to help you discount it correctly rather than to persuade you.
The test that matters most

We shuffled the signal 100 times and re-ran everything

The obvious objection to any backtest is that the machinery itself manufactures the result. So we destroyed the signal and kept the machinery: within each rebalance date the scores were randomly reassigned between companies — preserving the distribution, the missing data and the correlation between themes, destroying only the link between a company and its own score — and the whole pipeline was run again, 100 times, on seeds 1000..1099.

Grey is where those 100 deliberately worthless runs landed. The dark line is the real result.

Top-decile alpha, per quarter beat all 100 noise runs
real 7.17%
noise bar
Typical noise run 0.25%  ·  best of 100 noise runs 2.78%  ·  real 7.17%
Top-decile alpha, t-statistic beat all 100 noise runs
real 4.38
noise bar
Typical noise run 0.31  ·  best of 100 noise runs 3.32  ·  real 4.38
Long-short spread, t-statistic 3 of 100 noise runs matched or beat it
real 2.62
noise bar
Typical noise run 0.12  ·  best of 100 noise runs 3.78  ·  real 2.62
Decile ladder ordering 1 of 100 noise runs matched or beat it
real -0.89
noise bar
Typical noise run -0.07  ·  best of 100 noise runs -0.89  ·  real -0.89  ·  lower is better here: −1.0 is a perfectly ordered ladder

2 of these beat every single noise run outright. The rest did not, and the counts are on the labels above rather than rounded away: long-short spread, t-statistic was matched or beaten by 3 of 100, decile ladder ordering was matched or beaten by 1 of 100. That is why the long-only ranking is the claim this site leads with and the long-short spread is not.

Counts derived at page load from PLACEBO_HAC.json, which retains all 100 draws (pre-registered as PREREG_ma19_ma13_recalibration.md). The "noise bar" is the 95th percentile of those draws. A tie counts against us: a noise run that merely equals the real result is counted as having matched it.
Against what you could have bought instead

Three benchmarks, including the one that flatters us least

The top-tenth ranked by our score, rebalanced quarterly, held for one quarter, equally weighted. Measured against three different comparisons rather than the single most favourable one.

BenchmarkIt returnedTop decileDifferencetQuarters ahead
equal-weight universe (incumbent)
uninvestable — you cannot buy this
every name in the panel, cost-free — uninvestable, the number every historical alpha figure in this project used
18.1% 25.3% +7.17% 4.38 71%
cap-weighted panel average
closest investable analogue buildable from the panel itself
14.9% 25.3% +10.46% 4.29 68%
SPY total return over the same windows
what the user's obvious alternative actually returned
15.3% 25.3% +9.99% 3.77 64%

All figures annualised and gross of trading costs. The equal-weight universe is listed first because every historical figure in this project was measured against it — and it is the hardest of the three, because an equally weighted basket of every company in the panel, charged nothing to trade, beat the S&P over this window. The SPY row is the one that answers "versus just buying the index".

Does the score actually sort?

The full ladder, not just the top

A top decile can look good by luck. The question is whether the whole ranking orders returns. Annualised return by score decile, best-ranked first.

Top 10%
25.3%
Decile 2
20.4%
Decile 3
19.5%
Decile 4
17.5%
Decile 5
17.3%
Decile 6
17.8%
Decile 7
18.2%
Decile 8
16.3%
Decile 9
14.6%
Bottom 10%
14.3%

It sorts, and imperfectly — deciles 4 through 7 are effectively tied and one is out of order. The ordering statistic is -0.891, where −1.0 would be a perfect ladder. What carries the result is the top decile and the bottom two; the middle is noise.

What the average hides

It lost to the market in 29% of quarters

An annualised average is one number standing in for 69 quarters, and quoting it alone is how a strategy that is often behind gets described as though it were always ahead.

Quarters behind the benchmark
20 of 69
Nearly three in ten. If you held this you would have spent a lot of time wondering whether it had stopped working.
Worst single quarter
-6.83%
Relative to the benchmark, in 2016-01-20. The best was +11.47% in 2022-07-22.
Typical quarter (median)
+1.41%
Below the mean of +1.79%, because the distribution is skewed: a few strong quarters pull the average above the typical one.
Worst peak-to-trough, after costs
-28.5%
Driven by one quarter — the Covid crash. A ranking of companies cannot protect you from the market falling as a whole.
Trading costs

How expensive would trading have to get before this stops working?

"Net of costs" asks you to believe our cost estimate. The breakeven does not: it is the cost at which the edge reaches exactly zero, so you can compare it to whatever you think you would really pay.

Breakeven trading cost
134 bps
Per trade, each way. Above this the edge is gone.
Estimated actual cost
33 bps
A margin of about 4.0×. The book turns over 261% a year, so costs are not a rounding error.

After costs the top decile's edge over the equal-weight universe falls from 8.12% to 6.07% a year.

What the cost model does not capture

  • keyed on point-in-time market cap ONLY — no spread, no average daily volume, no price level, no participation rate
  • the equal-weight benchmark is charged ZERO cost while the strategy pays, so every 'alpha versus equal-weight' figure compares a book that trades against one that does not
The part most sites leave out

The bars it fails

Four standard thresholds. It clears one and fails three. All four are here because a page that showed only the passing one would be advertising, not evidence.

ThresholdResultBar
Long-short t vs. this project's own placebo floor
The 95th percentile of 100 shuffled-signal runs. Beating it means the result is bigger than 95 out of 100 runs on a signal known to be worthless.
2.62 2.284 clears
Long-short t vs. the Harvey-Liu-Zhu multiple-testing hurdle
A bar that rises with the number of tests you have run. We have run a lot. This is the headline's clearest failure and the artifact records both sides of the argument.
CLEARS the bar measured against the project's own placebo and FAILS the bar derived from counting its own trials. Neither is 'the' answer. HLZ prices the best of N draws; the deployed composite is flat 1/7, never tuned, and cpcv.adopt is false on every run, so the logged trials are overwhelmingly REJECTED ALTERNATIVES to it rather than candidates it beat.
2.62 3.29 fails
Probability of backtest overfitting
Fails. And the bar is close to useless here: on a signal shuffled into pure noise this statistic reads about 0.47 on average, so roughly half of all worthless signals 'pass' it. We report it failing rather than quietly dropping a measure that does not flatter us.
0.733 0.5 max fails
Deflated Sharpe ratio
Fails the conventional 0.95 bar. It is a genuine deflated figure — it is charged every one of the tests below, not the eight a naive version would use — and it sits above all 100 shuffled runs. Both halves are true and we do not quote one without the other.
0.786 0.95 fails
What we tested and threw away

558 recorded tests. The overwhelming majority failed.

The reason a backtest is weak evidence is that you can keep trying until something works. The defence is to write down every attempt, including the ones that failed, so the successes can be discounted by how many shots were taken. Every test below was registered in writing — hypothesis, threshold and direction — and committed to version control before it was run.

Stock-ranking tests
248
Each one raises the bar the headline has to clear. That is why the multiple-testing hurdle above is failed: we counted honestly.
Options tests
310
Which produced no usable signal at all. See below.

Ideas that sounded good and were rejected on measurement

  • Sector-neutral ranking. Ranking companies against their own sector instead of the whole market. Tested three times, on two different panels, rejected every time. Now permanently closed.
  • A machine-learning model over the same inputs. A gradient-boosted tree trained on one half of the history. Out of sample its deciles ran backwards — its top decile underperformed its bottom one. The plain equally-weighted sum beat it in both directions of the split.
  • Tuning the weights. Five separate schemes for weighting the themes — shrinkage, inverse-volatility, recency-weighting and two others. All five rejected. The shipped model weights every theme equally and has never been tuned, because tuning it on this data made it worse.
  • The options entry signal. Our own options alert was tested against simply picking random days, and lost. It is not sold, not recommended, and the finding is published rather than buried.
  • Using the stock score to pick options. Rejected in both the versions we tried. The score describes the underlying company, not the trade.
  • A dozen published academic anomalies. Short-term reversal, idiosyncratic volatility, maximum daily return, low volatility, analyst estimate revisions, return seasonality and others. None replicated here at a threshold set in advance.

And corrections we found in our own work

  • An early version of the panel had an inverted universe in its first third. Fixing it cut the headline alpha by roughly a third. The corrected figure is the one on this page.
  • Five factors were silently empty for the project's entire history — they computed to nothing and raised no error. A coverage guard now fails the run instead.
  • Option strike prices were being compared against split-adjusted share prices, which manufactured a fake +31,000% trade. Repaired at source; no published verdict changed.

None of these were discovered by an outside auditor. They were found, written down and published by the project itself, which is the only reason you are reading about them.

Don't take our word for it

How to check this yourself

Four things you can verify without any special access, in rough order of how much effort they take.

1. Check that the numbers on this page match the artifact

Every figure above is read at request time from BACKTEST_RESULTS.json and MA19_RECALIBRATION.json. Nothing is typed into this page. The run is stamped 57bc3f2, dated 2026-08-14 — if that date is old, these numbers are old, and this page will say so rather than quietly refresh the label.

2. Check the arithmetic against a benchmark you can price yourself

The SPY row above says the S&P returned 15.3% a year over 2009-01-15 to 2026-01-28. That is a public number you can look up in thirty seconds. If our benchmark is wrong, everything measured against it is wrong, and it is the easiest thing on this page to falsify.

3. Ask the question that breaks most backtests

The usual killer is look-ahead: scoring a company in 2012 using a number that was not published until 2013. This panel is rebuilt as of each date from filings with their real filing dates, prices with no future adjustment, and institutional holdings lagged by their actual disclosure deadline — and companies that went bankrupt or were delisted stay in, so the record is not a survey of survivors. The How it works page explains the construction and the specific places it is still imperfect.

4. Read the tests that failed

Each one was registered before it ran, with the threshold written down first, and the verdict recorded whichever way it came out. That record is what makes the 248 figure above meaningful rather than decorative — a count of attempts is only honest if the failures were logged at the time rather than reconstructed afterwards.

The strongest evidence is not on this page

The most convincing test this project has run is an out-of-sample replication on international data: the same untuned model, mapped onto a different vendor's dataset, different countries, a different construction and a different period. It works there, and the United States is the weakest region tested. Those figures are derived from a dataset licensed for non-commercial research only, so they cannot be published on a product page. That is a licensing constraint, not a hedge — and a page claiming to show proof should say what it is leaving out and why.

Read this before you use any of it

What this is not

  • It is not a track record. No money was managed. Backtested results are hypothetical and benefit from knowing which companies still existed to be studied.
  • It is not evidence that it will keep working. The result comes from one panel over one period. Published factor premia routinely shrink after they are published, and there is no reason to assume this one is exempt.
  • It is not advice, and it is not tailored to you. A score is a model output about a company, produced without any knowledge of your circumstances, your tax position or your time horizon.
  • It is not a reason to concentrate. The figures above describe a basket of about 156 companies — the top tenth of a typical cross-section — rebalanced quarterly and equally weighted. Nothing here supports buying a handful of names, and a single name can go to zero regardless of its score.
  • The middle of the ranking means very little. Deciles four through seven are indistinguishable from each other. Treating a score of 61 as better than a 57 reads precision into the number that is not there.