We tested our own product. Here is everything it produced.

Methodology · the audit

Every figure below comes from committed, seeded artifacts (run 2026-08-01). Re-running the scripts against the same data reproduces the files byte-for-byte. The unfavourable numbers get the same type size as the favourable ones. Anyone can claim 70%; nobody publishes their frozen model returning zero.

Protocol

Section A · What we tested and how

Contract signals versus realized sector-relative alpha over 7/20/60/90 trading days (α = the stock's return minus its sector ETF). A linear model on the 9 score components was fit on 2021–2023 only (n=646), weights frozen with the engine's own postprocess, scaling bounds included, then scored on 2024–2026 (n=618) untouched by the fit. Outcomes winsorized at ±150% at the write site. The script is committed as scripts/walk-forward.ts.

The frozen model has no alpha.

Section B · Result 1

Out-of-sample Spearman IC +0.003, statistically nothing. A linear model on these 9 components, frozen in December 2023, would not have ranked anything usefully in 2024–2026. In-sample its top quintile hit 57%; none of it transferred.

Cellmean (w)mediansdp10hitn
Q1-2.5%-3.6%27%-23.2%40%124
Q2+1.8%-1.2%40%-32.3%49%124
Q3+4.2%-0.3%41%-28.1%49%123
Q4+6.0%-2.7%59%-56.2%48%124
Q5 (top)+2.3%-5.1%39%-27.3%45%123
top decile+7.8%-1.8%38%-19.4%49%61
ALL test+2.4%-2.7%42%-32.9%46%618

2024–26 test window · winsorized mean · interpolated median · p10 = worst decile · n in every cell

What survives out of sample: award size.

Section C · Result 2

Expanding-window cross-validation, genuinely-future fold (fit ≤2024 → 2025–26), investable universe (market cap ≥ $5M, n=284):

RankerICtop-quintile hitbase ratetop mediantop p10n
ridge_refit (frozen components)−0.12240%47%-8.6%-33.6%284
live ML score−0.06046%47%-5.8%-27.4%284
contract_value (award size)+0.19753%47%+3.1%-25.7%284

Full universe (n=321): contract_value IC +0.204, top hit 55% vs 47% base, top median +3.3%. Raw award size is the only ranker that survives, so the shipped score IS the award-size percentile, published openly. Against the base rate this tilt is roughly 2 sigma: real-looking, not bulletproof.

Ceiling and floor.

Section D · The caveat we cannot clear

The ML pipeline's training is not provably free of hindsight; we cannot certify it never saw the test years. Treat the live-score numbers as the ceiling and the frozen model's zero as the floor. We ship what sits between them and survives: award size.

Against 2,000 seeded random draws.

Section E · The universe baseline

2,000 seeded draws from our own signal ledger (1,160 outcomes), matched to the top rank's market-cap mix (top rank n=296). The top-rank median beat 98.0% of draws at 60 days (+0.98% vs draw median -2.16%) and 95.3% at 90 days, but only 55–58% at 7–20 days. The tilt lives at long horizons; at short horizons the top rank is indistinguishable from chance.

Market cap is not the driver.

Section F · The size factor is dead, both ways

Spearman of log-market-cap vs alpha: −0.010 (7d), +0.020 (20d), +0.080 (60d), +0.108 (90d). Size is not driving the separation, in either direction.

Even the top rank loses often.

Section G · The left tail

Top-rank dispersion is wide (sd ≈ 39% winsorized), and one in ten top-rank signals still loses -25.7% or more versus its sector over 60 trading days. Every hit rate we publish carries its worst decile alongside.

This is a binary discriminator.

Section H · The ranking admission

Below the top rank, tiers are statistically indistinguishable from each other. This is top rank versus everything else, not a multi-level ranking. We say it here before you find it in the grid on the performance page, where the full grid is published anyway.

The distribution is outlier-dominated.

Section I · Why we don't lead with means

Unrealizable microcap prints dominate raw averages, so outcomes are winsorized at ±150% at the write site, and we publish median, winsorized mean, and worst decile together. A mean quoted alone is marketing.

Coverage, latency, and a disclosed tilt.

Section J · What we therefore claim

Arcthane claims coverage of the public record, same-day latency, and a ranking with a measurable, disclosed tilt, with the base rate and n attached every time it is quoted. Not returns.

Walk-forward artifacts: engine v3.0.0 · draw baseline: v3.1.0 ledger · both run 2026-08-01 · winsor ±150%