Twenty-seven things we tested and threw away.
This product does not predict prices. That is not a stylistic choice or a legal hedge — it is what is left after two years of testing whether we could, writing the pass mark down first, and losing.
Every claim below was specified before the data was fitted: the hypothesis, the measurement, and the number that would count as success. None was retuned after the fact to rescue it. The page exists because a tool that tells you what it cannot measure is worth less than nothing unless you can check that it means it.
A real signal, killed for being redundant
The last test asked whether the structural block at the heart of this engine — how many genuinely independent bets a portfolio holds — explains how hard that portfolio falls in a crisis, over and above plain volatility, beta, CVaR and trailing drawdown. Sixteen portfolios, three crises, forty-eight observations. Measured sixty-three trading days before each crisis began, so nothing could see what was coming.
What was written down beforehand, and what came back
| Test | Needed | Measured | |
|---|---|---|---|
| Explains crisis drawdown beyond the classic metrics R² rises from 0.8805 to 0.8907 |
ΔR² ≥ 0.05 | +0.0103 | missed |
| More independent bets means a smaller fall | coefficient < 0 | −0.0151 | met |
| Ranks portfolios correctly inside each crisis | Spearman ≤ −0.30 | −0.750 | met |
Why the two passes were not enough
The ranking worked, and worked hard — inside every one of the three crises, the portfolios with fewer independent bets fell further.
And it still bought nothing, because volatility had already said it. Plain trailing volatility ranked the same portfolios at +0.781 against the same drawdowns. Two instruments pointing at one fact is one fact. Shipping the second as a discovery would have been the most flattering thing this project could have done to itself.
Forty-eight observations is a small panel and the result is reported as directional, not conclusive. The pre-registration says so too — it was written before the answer was known, which is the only time that sentence is worth anything.
Twenty-six before it
All of them asked a version of the same question: can something visible today tell you something useful about tomorrow, at the scale one person actually invests at? Each had its pass mark fixed in advance. Each was abandoned when the mark was missed, rather than adjusted until it cleared.
- A drawdown predictor from market-wide stresskilled The first version, and the one that set the pattern for every version after it.
- Stress that persists, and a threshold to act on itkilled
- Hidden-Markov regime detectionkilled The regimes were legible after the fact and useless before it.
- Conditional return distributions given the stress statekilled
- Quantile forecasting of the next movekilled
- Post-crash reversion — buying the bounce systematicallykilled
- Euphoria fade — selling the top of a runkilled
- Trend following, twice, on two different constructionskilled
- A phase-density measure borrowed from physicskilled Steel-manned deliberately before testing, so the kill could not be blamed on a weak version.
- A probability layer over the stress signalkilled Held up on one market in isolation and failed when pooled — which is the version that would have been shipped.
- Funding-rate carry as a systematic edgekilled
- Two new mechanisms, generated specifically to break the losing streakkilled
- A crash-test service for trading botskilled Not on statistics — on customer discovery. The people who had the problem already had a good-enough answer for free.
- Two marketplace and opportunity-scanning productskilled One was already being done well by an incumbent. Looking first cost a week; not looking would have cost a year.
One earlier result was killed twice: published internally with an effect size that a later audit found had been credited to the wrong cause. The corrected figure was a fraction of the original. It is on this list rather than quietly deleted because a project that only records other people's errors is not keeping a log, it is keeping a brochure.
Description, and the discipline that produced it
What none of those tests could kill is measurement of what is already there: how concentrated a portfolio actually is, what it holds through the funds it holds, how much of its result came from a currency rather than from an asset, what it did in each crisis it was old enough to live through. None of that requires knowing the future. All of it is checkable against your own price history.
The same rule, applied to our own optimiser
Fitting an allocation to its own history always looks good. So the app fits it on the years before each decision, holds it for twelve months, and refits — eighteen times over — then shows the result whichever way it falls.
Return per unit of risk, eighteen refits, 2009 to 2026. On this book the fitted allocation came out ahead. On other books in testing it lost to equal weight outright, and the app says so there too. It also rebuilt a fifth of the portfolio at every refit, which is its own kind of answer.
That is the whole offer. Not a forecast you have to trust, but arithmetic on your own holdings, with the window it was measured over printed next to every figure, and the things it cannot measure named instead of skipped.