We tested our own product. Here is everything it produced.
Methodology · the audit
Every figure below comes from committed, seeded artifacts (run 2026-08-01). Re-running the scripts against the same data reproduces the files byte-for-byte. The unfavourable numbers get the same type size as the favourable ones. Anyone can claim 70%; nobody publishes their frozen model returning zero.
Protocol
Section A · What we tested and how
Contract signals versus realized sector-relative alpha over 7/20/60/90 trading days (α = the stock's return minus its sector ETF). A linear model on the 9 score components was fit on 2021–2023 only (n=646), weights frozen with the engine's own postprocess, scaling bounds included, then scored on 2024–2026 (n=618) untouched by the fit. Outcomes winsorized at ±150% at the write site. The script is committed as scripts/walk-forward.ts.
The frozen model has no alpha.
Section B · Result 1
Out-of-sample Spearman IC +0.003, statistically nothing. A linear model on these 9 components, frozen in December 2023, would not have ranked anything usefully in 2024–2026. In-sample its top quintile hit 57%; none of it transferred.
| Cell | mean (w) | median | sd | p10 | hit | n |
|---|---|---|---|---|---|---|
| Q1 | -2.5% | -3.6% | 27% | -23.2% | 40% | 124 |
| Q2 | +1.8% | -1.2% | 40% | -32.3% | 49% | 124 |
| Q3 | +4.2% | -0.3% | 41% | -28.1% | 49% | 123 |
| Q4 | +6.0% | -2.7% | 59% | -56.2% | 48% | 124 |
| Q5 (top) | +2.3% | -5.1% | 39% | -27.3% | 45% | 123 |
| top decile | +7.8% | -1.8% | 38% | -19.4% | 49% | 61 |
| ALL test | +2.4% | -2.7% | 42% | -32.9% | 46% | 618 |
2024–26 test window · winsorized mean · interpolated median · p10 = worst decile · n in every cell
What survives out of sample: award size.
Section C · Result 2
Expanding-window cross-validation, genuinely-future fold (fit ≤2024 → 2025–26), investable universe (market cap ≥ $5M, n=284):
| Ranker | IC | top-quintile hit | base rate | top median | top p10 | n |
|---|---|---|---|---|---|---|
| ridge_refit (frozen components) | −0.122 | 40% | 47% | -8.6% | -33.6% | 284 |
| live ML score | −0.060 | 46% | 47% | -5.8% | -27.4% | 284 |
| contract_value (award size) | +0.197 | 53% | 47% | +3.1% | -25.7% | 284 |
Full universe (n=321): contract_value IC +0.204, top hit 55% vs 47% base, top median +3.3%. Raw award size is the only ranker that survives, so the shipped score IS the award-size percentile, published openly. Against the base rate this tilt is roughly 2 sigma: real-looking, not bulletproof.
Ceiling and floor.
Section D · The caveat we cannot clear
The ML pipeline's training is not provably free of hindsight; we cannot certify it never saw the test years. Treat the live-score numbers as the ceiling and the frozen model's zero as the floor. We ship what sits between them and survives: award size.
Against 2,000 seeded random draws.
Section E · The universe baseline
2,000 seeded draws from our own signal ledger (1,160 outcomes), matched to the top rank's market-cap mix (top rank n=296). The top-rank median beat 98.0% of draws at 60 days (+0.98% vs draw median -2.16%) and 95.3% at 90 days, but only 55–58% at 7–20 days. The tilt lives at long horizons; at short horizons the top rank is indistinguishable from chance.
Market cap is not the driver.
Section F · The size factor is dead, both ways
Spearman of log-market-cap vs alpha: −0.010 (7d), +0.020 (20d), +0.080 (60d), +0.108 (90d). Size is not driving the separation, in either direction.
Even the top rank loses often.
Section G · The left tail
Top-rank dispersion is wide (sd ≈ 39% winsorized), and one in ten top-rank signals still loses -25.7% or more versus its sector over 60 trading days. Every hit rate we publish carries its worst decile alongside.
This is a binary discriminator.
Section H · The ranking admission
Below the top rank, tiers are statistically indistinguishable from each other. This is top rank versus everything else, not a multi-level ranking. We say it here before you find it in the grid on the performance page, where the full grid is published anyway.
The distribution is outlier-dominated.
Section I · Why we don't lead with means
Unrealizable microcap prints dominate raw averages, so outcomes are winsorized at ±150% at the write site, and we publish median, winsorized mean, and worst decile together. A mean quoted alone is marketing.
Coverage, latency, and a disclosed tilt.
Section J · What we therefore claim
Arcthane claims coverage of the public record, same-day latency, and a ranking with a measurable, disclosed tilt, with the base rate and n attached every time it is quoted. Not returns.
Walk-forward artifacts: engine v3.0.0 · draw baseline: v3.1.0 ledger · both run 2026-08-01 · winsor ±150%