Mercury

An engine that manufactures edge — and proves it out of sample.

Mercury is a private, multi-strategy crypto trading engine for managed capital. Capital is distributed across distinct sleeves with independent theses, balanced long and short across pairs, and governed by deterministic risk logic at both position and portfolio level.

Not a strategy. An engine that generates, tests, promotes, and retires strategy.

3
sleeves cleared a pre-registered holdout
2,261
positions in the ranked cohort
73
symbols traded, HHI 0.02
+0.35 pp
portfolio mean, every position counted

Current observation window. Return per position, size-weighted, single-sleeve attribution, open risk marked to market. Full method and per-sleeve numbers in the evidence dossier.

Most trading systems are one bet wearing a uniform

One narrow setup, one regime, one founder's discretion — and a track record that is really a single sample. When the regime turns, there is nothing behind the first line.

Multi-sleeve by construction

Capital is spread across sleeves with separate theses, separate horizons, and separate failure modes. No single sleeve is load-bearing for the portfolio.

Neutral posture

Long and short books run concurrently across pairs. Leverage operates inside a neutralised structure, not as a directional multiplier.

Attribution, not aggregate

Every closed position carries the reason it closed. Performance is decomposed by sleeve, side, and exit path — so a good average can never hide a broken component.

Breadth over concentration

The current cohort spans 73 symbols with a Herfindahl index of 0.02. The result is a distribution, not two lucky tickers.

The research loop is the asset

Any individual edge decays. What compounds is the rate at which an operation can form a hypothesis, specify it precisely, deploy it safely, and kill it honestly. Mercury is built as that loop, with AI embedded where it multiplies throughput and deliberately excluded where determinism is worth more.

  1. 01

    Hypothesis, written before code

    Every sleeve starts as a written thesis: what market state carries the edge, what must stay true for it to remain itself, and which nearby state is explicitly not it. The thesis is versioned and reviewable before a single order is designed.

  2. 02

    Executable specification

    The thesis compiles into a machine-checkable specification — required facts, admission conditions, invalidation conditions, exit geometry. Two independent implementers must derive the same verdicts from it, or the spec is rejected as ambiguous.

  3. 03

    Deterministic engine

    Signal evaluation, risk sizing, and exit management are deterministic code paths, not model outputs. AI accelerates hypothesis generation and development throughput; it does not sit in the order path.

  4. 04

    Three parallel tracks

    The engine runs as three independently deployed instances, each on its own configuration pin and its own datastore. A sleeve that only prints on one track is treated as a configuration artifact until it replicates.

  5. 05

    Pre-registered holdout

    The holdout cutoff is locked before any sleeve is scanned. A sleeve is promoted only if the out-of-sample half stands on its own — not because the full-period average looks good.

  6. 06

    Promotion and demotion

    Every promoted sleeve ships with a written falsification condition and a review window. Sleeves that fail it are demoted on schedule, not defended.

Where this goes next

Athanor — taking the human out of the middle of the loop

The next iteration turns the loop above into a closed-loop system that selects its own next experiment against a real research budget, records what it ruled out as carefully as what it found, and rebuilds its own instruments when a question is not yet answerable. The architecture is transferred from self-driving laboratories in protein engineering.

Read the Athanor design →

Evidence, on the terms a sceptic would set

The observation window, the holdout cutoff, and the mark-to-market timestamp were all fixed before any sleeve was ranked. Positions are counted from the date they were opened, not the date they closed, so fast winners cannot crowd the sample. Open risk is marked to market and carried into the result rather than excluded.

SleevenMeanDiscovery → holdout
A — short book, tactical158+0.594+0.596 → +0.592
B — long book, regime84+0.751+0.862 → +0.718
C — short book, structural78 / 72+0.782 / +0.520replicates on two tracks

Mean return per position in percentage points of notional, size-weighted. Confidence intervals, portfolio totals, exit attribution, sleeves that were not promoted, and the written falsification conditions are in the dossier.

Risk and governance

  • Deterministic order path. Sizing, stop placement, and exit management are explicit code with test coverage. No model output reaches the venue unmediated.
  • Fail loud, not quiet. Broken internal state raises an error and halts the affected path. Silent fallbacks that keep a process alive while suppressing trades are treated as defects, not safety.
  • Position and portfolio limits. Risk is bounded per position, per sleeve, and per book, with concentration monitored as a first-class metric.
  • Full decision trace. Every admission, rejection, and exit is reconstructable from logs. Post-trade review can answer why, not just how much.
  • Written kill criteria. Each promoted sleeve carries the condition under which it will be demoted, fixed in advance of the review window.

Capital fit

For selective partners who want access to a private, evolving engine with disciplined risk and real technical depth behind it. The value on offer is not one month's return — it is ownership-like access to the loop that produces the next sleeve after this one stops working.

The conservative operating profile I am prepared to defend — target return, leverage envelope, and maximum drawdown — is shared directly in an allocation conversation, under the specific mandate being discussed, rather than published as a headline number.

Next step

I am opening conversations around an initial pilot allocation with aligned capital partners. If the fit is right, the next step is a direct discussion on mandate structure, operating model, and pilot terms — [email protected].

All trading involves risk of loss. Any allocation discussion is subject to mandate design, operational review, and live risk controls.

kaido.team — one operator, a fleet of agents, under one flag.

the name