Mercury · evidence dossier

What survived a holdout that was locked before we looked.

A track record is only worth reading if the rules were fixed before the results were known. This page states the rules first, then the numbers, then the sleeves that did not make it, then the conditions under which the ones that did will be demoted.

← Back to the engine overview

Method, stated first

Three windows were fixed before any sleeve was scanned: the observation window, the discovery/holdout cutoff, and the timestamp at which open positions are marked. Nothing below was re-cut after the fact.

Cohort
Every position opened inside the current observation window. Membership is decided by entry date, so fast closes cannot crowd the sample.
Holdout
The cutoff was locked before the first sleeve was scanned. Everything before it is discovery; everything after it is out of sample.
Metric
Return per position in percentage points of notional, size-weighted within the position. Closed positions use realised return; open positions are marked at the cutoff price.
Grain
Single-sleeve positions only. Positions carrying overlapping sleeve exposure are excluded from ranking, so no sleeve is credited with another's result.

Three tracks are three independently deployed instances of the engine, each on its own configuration pin and its own datastore. A result that appears on one track is a candidate. A result that replicates across tracks is an edge.

Discovery versus holdout

The only chart that matters here. A sleeve is interesting when the out-of-sample bar stands close to the in-sample bar; a tall discovery bar next to a short holdout bar is a warning, and is reported as such.

Discovery Holdout (out of sample) scale 0 → 1.2 pp
  • Sleeve A — short book, tactical

    Track II

    0.596
    0.592
  • Sleeve B — long book, regime

    Track I

    0.862
    0.718
  • Sleeve C — short book, structural

    Track III

    1.126
    0.336
  • Sleeve C — short book, structural

    Track I

    0.730
    0.400

Sleeve A holds its out-of-sample mean almost exactly. Sleeve B gives back roughly a sixth. Sleeve C decays hard on Track III and mildly on Track I — which is why it is presented as a two-track result rather than a Track III headline.

Promoted sleeves

Ranked on expectancy, sample size, interval width, holdout survival, and symbol breadth. Pockets with fewer than 20 positions, or with a failing holdout, are ineligible regardless of average.

SleeveTracknMean pp95% CISymbolsHHI
A — short book, tacticalII158+0.594[+0.34, +0.85]730.02
B — long book, regimeI84+0.751[+0.32, +1.21]730.02
C — short book, structuralIII78+0.782[+0.30, +1.33]540.02
C — short book, structuralI72+0.520[+0.06, +0.93]540.02

Strongest current position

Sleeve A on Track II. Out-of-sample mean is within 0.004 pp of in-sample — the closest thing to a non-result from the overfitting test that this dataset can produce. It clears on 158 positions across 73 symbols with a concentration index of 0.02, and the same sleeve prints positive on Track III (+0.367, n=115). On Track I it is still positive (+0.304) but its interval touches zero, so Track I is reported as unconfirmed rather than supporting.

Portfolio level, all sleeves

Every position in the cohort, promoted sleeves and rejected ones together. This is the number that describes the engine as deployed, and it is deliberately lower than the crowned sleeves — the difference is the cost of running research in production.

Track I

+0.35 pp

Positions
651
95% CI
[+0.22, +0.48]
Open, marked
16

Track II

+0.33 pp

Positions
856
95% CI
[+0.24, +0.42]
Open, marked
5

Track III

+0.39 pp

Positions
754
95% CI
[+0.28, +0.50]
Open, marked
24

All three intervals exclude zero. Settled by close date instead, the same period books 857, 1,095, and 931 positions for a summed +224, +282, and +284 pp — the larger, more flattering number, and precisely the one this study refuses to rank on.

Where the money is actually made and lost

Each closed position records the path it exited through. Decomposing by exit path is how a sleeve's average is audited: it separates the mechanism that earns from the mechanism that bleeds, and it points at the next experiment rather than the next slide.

Sleeve A, Track II — both books

BookExit pathnMean
Short bookTrailing stop (profit harvest)122+0.95
Short bookThesis invalidation34-0.67
Long bookTrailing stop125+0.56
Long bookThesis invalidation42-0.69

Invalidation costs both books the same. The split between them is entirely in the exit geometry of the winners — which localises the edge to a design decision we control, not to a market accident.

Sleeve B, long book — Track I versus Track II

TrackExit pathnMean
ITarget reached15+2.70
ITrailing stop61+0.28
IITrailing stop84+0.80
IIThesis invalidation75-0.52

Same sleeve, two parameterisations. Track I lets a small tail of positions reach target and pays +2.70 pp on them; Track II invalidates roughly 45% of its closes and taxes the average away. The open work is to bring Track II's invalidation behaviour to Track I's, not to switch the sleeve off.

Cleared the bar, not promoted

These pockets passed the holdout and the interval test and were still held back. Publishing them is the point: a shortlist with no rejects behind it is a shortlist that was chosen after the fact.

PocketnMeanWhy it was held back
Sleeve D — long book, swing (Track III)25+0.85Sample too thin to carry a promotion: only 9 positions fall in the discovery half.
Sleeve B — short book (Track III)53+0.57Weaker than the long side of the same sleeve, and does not replicate on Track I, where its holdout fails.
Sleeve C — long book (Track II)48+0.47Real, but a smaller sample and the weaker sibling of a short book that replicates on two tracks.
Sleeve A — long book (Tracks II, III)170 / 184+0.25 / +0.22Positive and stable, but materially weaker than the short book on the same engine. Held, not crowned.

What this study refused to do

  • No sleeve was ranked on its full-period average alone; the holdout half had to stand by itself.
  • No holdout cutoff was chosen after seeing results — the date was fixed before the first scan.
  • No expectancy was computed on positions grouped by close date, which over-represents fast winners.
  • No sleeve was crowned because it beat a sibling; a contrast is mechanism evidence, never a ranking.
  • No positions were dropped for still being open — open risk is marked to market and carried into the mean.
  • No sleeve was retired on a single weak window without its written falsification condition being met.

How each result dies

Written before the next window opens, so a demotion is a scheduled outcome rather than an argument.

Sleeve A, Track II
On the next seven-day window, the holdout mean is at or below zero, or the confidence interval covers zero. Demote.
Sleeve B, Track I
The holdout mean turns non-positive on the next window, or the profit tail collapses onto a single symbol. Demote.
Sleeve C, Track I
The confidence interval covers zero on the next holdout at a sample of 40 or more. Demote.

Next step

Position-level data, the full sleeve inventory including everything that failed, and the operating profile open up in a direct conversation — [email protected].

All trading involves risk of loss. Past results do not predict future performance.

kaido.team — one operator, a fleet of agents, under one flag.

the name