How it works
What the model does, how it was tested, and — the part
most tools leave out — where it is weak.
Educational research tool — not investment advice, and not a recommendation to buy or sell any security. Backtested results are hypothetical, come from one 18-year dataset the model was also tuned on, and are not a promise about the future. The forward track is a model portfolio and a sandbox paper account — no money is invested in either, so no figure here is a return anyone received — and it is short. You can lose money. Do your own research.
The short version
After each market close the ~800 most liquid US common stocks are scored on the model's
weighted themes — value, quality (growth, for younger companies), momentum, size, capital
discipline, institutional positioning and insider activity, each carrying an equal share.
Each theme is built from individual numbers, each number is standardised across the names
scored that day, and the themes are combined into one composite. The Hot Stocks tab is
that ranking. Two of the seven — institutional and insider — have no live data source yet, so
today's live ranking runs on the other five while the backtest used all seven; the Hot Stocks
tab flags this on every scan.
The research also constructs a book from that ranking — the top decile of the large-cap
tier, score-weighted, capped at 8% per name — and tracks it forward against the S&P 500.
Both are shown in the app: the Valquo Index tab lists the book and the Track
Record tab its forward record. They are a model portfolio with no money in it,
published so the method can be checked, not as a recommendation to hold any of it.
Point-in-time, and why it matters more than the score
The backtest is built from a point-in-time fundamentals panel: on any historical date the
model sees only what was actually public and filed by then. Restated figures do not leak
backwards. Institutional (13F) holdings are lagged by their real filing deadline — the
panel's effective lag is about 111 days, more conservative than the 45-day rule. We checked
this the direct way: feeding the model fresher, not-yet-filed holdings makes it
weaker, not stronger, which is the opposite of what a look-ahead artefact does.
Survivorship
Delisted and acquired companies stay in the universe up to the date they disappear, using
a delisting mask rather than today's surviving-company list. A backtest run on the names that
still exist is a backtest of not going bankrupt. The live forward track has the same property
by construction: the picks are dated when they are made and measured forward, so nothing can
be quietly dropped after the fact.
Costs — quote the breakeven, not the net number
Trading costs are modelled, not assumed away. The top decile's breakeven cost is about
134 basis points one-way against a measured 33 bps real-world cost profile, on
roughly 261% annual turnover. We lead with the breakeven deliberately: it requires you to
believe no particular cost estimate. Borrow cost is not modelled, which affects the
long/short statistic but not the long-only book that the Index actually is.
How we decide something is real
- Combinatorial Purged Cross-Validation is the authority on weights, reporting the
probability of backtest overfitting (PBO) and a Deflated Sharpe that penalises how many
variants were tried. If it rejects, the shipped defaults stay — we do not quietly fall
back to the friendlier single-path result.
- Held-out confirmation. A change decided on one half of history has to survive
measurement on the half that did not inform it, in both directions. Several plausible
improvements were rejected exactly here.
- Full universe only. Verdicts come from the whole ~2,531-name panel. Smaller
samples systematically flatter results and are used only as smoke tests.
The one result that had to clear a threshold written down in advance
The composite is roughly one-seventh each of value, quality, momentum, size, investment
discipline and institutional ownership — nearly the standard factor set — so "it beat the
average" could not distinguish a real edge from an assembly of known premia. Two write-ups
were prepared before the regression ran: one to publish if the intercept cleared
Newey–West t > 2, and one saying the honest description was efficient factor
exposure if it did not.
It cleared it. Against the Fama–French five factors plus momentum, the top-decile spread
carries an intercept of +6.99%/yr (t = 3.98) over 68 non-overlapping windows,
2009–2025, with all six pre-registered specifications positive at t > 2 and spanning
+5.1% to +10.9%. Run through the same pipeline, a passive large-cap ETF returns
+0.68%/yr (t 1.58) — the placebo that shows the machinery reports nothing
significant when there is nothing there.
Read that as a research finding and nothing else.
It is measured on a
historical simulation, not an account. It is
not an expected
return, not an achievable return, and not a return anyone earned. It is one panel; the
t-statistic is not corrected for the equity trials logged behind it (248 by 2026-09-29, and
the count keeps rising — the live figure is on the
Proof page); the panel's
known-contaminated early period has been removed from the sample outright rather than
discounted, and the conservative single figure is the first half's
+5.19% if you
want one number. And it does not mean the strategy beats the
factor ETFs you can actually buy — that was tested separately and
was not demonstrated
(a +9.2pp margin at t 1.10, negative in the first half: a null).
How long the edge lasted, and what that does not license
Every headline figure this project publishes is measured over a 63-trading-day
forward window, which was an inherited default rather than a measured optimum. So the
composite was scored at 8 horizons — one quarter through two
years — from a single panel build, on the set of dates observable at every horizon, so
that the length of the window is the only thing that varies. The result:
CONSTANT-RATE.
In the backtest, the top decile of the hot list beat the equal-weighted universe by about 6.6% annualized over the next three months — and was still ahead by about 5.1% annualized two years later — even though a given name typically stays in the top decile for only one quarterly rebalance.
An independent route reaches the same place without touching the decile machinery: the
median per-date rank correlation between score and forward return rises with horizon,
0.0336 at one quarter to
0.0655 at two years. And the alpha is well
measured at every horizon — its t-statistic never falls below 3.16, and is 3.83 at two years
— so this is not a signal that survives because its error bars widened.
Three limits, and the third is the one people get
wrong. Measured on the corrected 2,531-name / 69-date panel: long-only top decile versus the equal-weighted universe; gross of costs; and it is the same single in-sample panel every other published figure comes from — not a forward test.
Second, a longer forward window is not new data: the eight horizons are eight views of
one sample, not eight samples, and the windows overlap almost entirely.
Third: It is not a finding that the book should rebalance less often — the list is re-ranked every quarter, and what was measured is how one quarter's selection went on to do, not a comparison of holding policies. What was measured is the buy-and-hold return of
the names picked on a single date; a quarterly-rebalanced list re-picks and compounds fresh
selections, and those are different claims. Only the first was tested.
The options side wins about a third of the time, and that is the design
The research also runs an options book, and its hit rate is 35-37%
depending on which book you measure. That number is the one most likely to be misread, so here
is the whole distribution rather than the average. Over
3,885 simulated trades on 187 names,
2016-01 to 2025-10:
- 1.4% — lost almost everything (worse than -90%)
- 58.2% — hit the stop (-45% to -90%)
- 5.0% — small loss (0 to -45%)
- 10.3% — small win (up to +100%)
- 25.0% — at least doubled (+100% or better)
The middle trade loses 52% of the
premium. The trades that at least doubled are 87%
of everything the winners made. A book like this is supposed to lose most of the time,
and a string of losses is its normal texture rather than a sign it has stopped working.
How long is an ordinary bad run? Expect losing streaks. Over 20 trades the typical worst run is 5 in a row, 44% of stretches contain a run of 6 or worse, and the record's worst at this scale is 20. Losing runs are
longer than a coin-flip model predicts, because trades opened near each other in time share a
market: the clustering measures 2.667 against a shuffled null
whose 95th percentile is 1.244
(1,000 shuffles, p < 0.001). Assuming
independence would put the 95th-percentile worst run at
10 instead of the measured 12 — so the tidy
arithmetic is the one that would cry wolf.
None of that says the options alerts work — they
were tested and they do not. Measured against random entry on the same names and dates, the
alert's choice of day subtracted value: −5.06 percentage points per trade, paired sign
test p < 0.00001. We publish the payoff shape so a losing streak is legible, not
as evidence of an edge. There is none demonstrated here, and the alerts are an idea generator
rather than a signal to act on.
Streaks measured on the 6,032-trade corrected-era random-entry control (the real book's own per-trade sequence is not banked); the control hits 37.2% against the book's 35.3%, so these runs are if anything too SHORT.
Where it is weak — read this part
- The ranking is coarser than its two-decimal scores look. We tested the score
itself against universes assembled by chance — same coverage, same weights, cross-theme
agreement destroyed — under a rule written down before the run. The per-name result
failed: On recent cross-sections the top decile as a group scores better than a chance-assembled book. Where an individual name sits inside that decile is not distinguishable from chance — and that second half holds on 45 of 69 dates tested. The group-level half holds on 21 of 69 dates tested — it is a property of recent cross-sections, not a standing one.
A second finding from the same test is worth carrying with it:
A high score can partly reflect missing data: names near the top are scored on less information than the average name, more so than chance would produce. What may no longer be said, on this evidence:
that a specific rank, or the gap between #3 and #12, means anything. This is a statement about
the ranking's precision, not about the backtested return spread — those are
different objects, and this test settles only the first.
- It has seen one dataset. Every backtested figure comes from a single 18-year
panel that the model was also tuned on. That is the fundamental limit, and no amount of
internal validation removes it. The live forward track is the only cure, and it is new.
- The Deflated Sharpe fails its conventional bar. Charged for every variant this
project has logged rather than the eight weightings it once counted, the backtest's
Deflated Sharpe is about 0.79 in the latest run — below the usual 0.95, and it falls
as more trials are logged. It does sit above the 95th percentile of 100 shuffled-signal
placebo runs, so it is distinguishable from noise without clearing the textbook threshold. An earlier version of
this page called the statistic undeflated: with only eight near-identical variants in
its denominator it saturated at >99.9%, deflating nothing. That was fixed by wiring in the
real trial count, and the page now reports the result that fix produced.
- Some themes are dormant. Sentiment (estimate revisions) has no point-in-time
source we trust and carries no weight. Low-risk carries none either: with both of its inputs
populated it measured indistinguishable from zero and was switched off.
- The live ranking runs on five of the backtest's seven themes. The scoring code is
the same function in both places — a replay of real historical dates through the live and
the backtest paths agrees to the last decimal — but institutional and insider positioning
have no live data source yet, so they contribute nothing to today's scores and their weight
is shared among the other five. The five-theme ranking has not been backtested on its own,
so its record is not the seven-theme one described here.
- Concentrated books are noisy. The 25-name construction scores higher in the
backtest than the decile, but it is the least reliable number in the study — which is
why breadth is the default.
- Live data has gaps. Live fundamentals come from a broker feed plus free public
sources rather than the licensed point-in-time export the backtest used, and a few fields
are missing for some names; the Hot Stocks tab says which source each scan used. Every page
dates itself, and a stale scan says so rather than being served as today's.
What this is not
It is not a forecast, not personalised advice, and not a claim that the next twelve months
will resemble the backtest. It is a disciplined, auditable ranking with its assumptions and
its failures written down. If a page here ever shows you a number without telling you
whether it is backtested or live, that is a bug — please report it.
It is also not a signal service. The app does show a model portfolio, options
alerts and forward records — but every one of them is a paper record: a model book
priced at closing marks, and a broker sandbox account with no real money. They are
published so the method can be checked against what actually happened next, and they are
labelled thin until they are long enough to mean anything. None of it is a return anyone
received, and the options alerts in particular were tested and found to lose to random entry
(above). Treat everything here as analysis you can check, not something to act on.
← Back to the app