Methodology
Docs · methodology
Two things are documented here: how a signal gets its score, and how we measure whether the scores mean anything. The second part is the one most tools skip. The full audited protocol and every number it produced, including the ones that lost, is published at arcthane.com/methodology.
How a signal is scored
The engine re-reads the public exhaust of federal money every two hours: awards, subawards, solicitations, budget lines, filings, FDA decisions, the Federal Register, patents, sanctions lists, plus market tape (options, sentiment) to detect what's already believed. Each event that maps to a public ticker becomes a candidate signal.
The contract score, 0–100, is deliberately simple and fully published: it is the percentile of the award's size within the historical scored population. Larger than 80% of scored federal awards → score 80, top rank. That's it. We ship the simplest ranker because it is the only one that survives out-of-sample testing on our own ledger; the machine-learning layer we also built does not, and we say so rather than sell it. Context components (materiality against revenue, sector and agency history, macro regime) are shown on every signal's detail page as diagnostics derived from the outcome ledger, not hidden weights. The transparency is the moat.
Timing is scored separately from strength. A large award read late gets a high score and an honest "likely priced in" flag; conflating the two is how tools trick users into chasing stale prints.
How we measure ourselves
When a signal fires, we record the stock's price and its sector ETF's price. Sixty trading days later, the difference in returns is the signal's alpha: sector-relative, so a rising market doesn't flatter the model and a falling one doesn't punish it.
- Median first, tail alongside. The median is the outcome a typical signal delivered; we publish the mean and the worst decile next to it, because medians alone hide the left tail.
- Fixed buckets, assigned at signal time, with no re-labeling after the outcome is known.
- Losing rows published. The track record shows every bucket, including the ones that underperform, starting with the unfiltered universe baseline, which is negative.
- Outcomes accrue at the horizon. Recent signals don't count until their window closes; there is no partial credit for a good first week.
- Winsorized at ±150%. Sub-$5M shells and unadjusted prints would otherwise dominate every mean; the clamp is applied when the outcome is written, not when it's convenient.
- Early sources are labeled. A signal source without at least 100 realized 60- and 90-day outcomes is marked 'early, tracking openly': a hypothesis, not evidence.
The baseline, and out of sample
Two tests stand behind every claim we make. First, the universe baseline: thousands of seeded random draws from our own signal ledger, matched to the top rank's market-cap mix, answer whether the top rank is actually better than chance, or just riding the universe's drift. Second, the out-of-sample protocol: rankers are frozen on data through a dated cutoff and scored on the years after it, with expanding-window folds keyed on the award date. Both are committed, seeded, reproducible artifacts, not a notebook we can quietly re-run, and the ranker only changes when a candidate beats the shipped one through that gate.
What we refuse to claim
Arcthane is research infrastructure, not investment advice, and the numbers say what they say: out of sample, top-ranked signals beat their sector modestly more often than the base rate, with wide dispersion, which means a large minority of top-ranked signals lose. We publish that instead of hiding it, because a track record you can't audit is marketing. The current figures are always live on the track record page.
- We never quote means without medians and the worst decile beside them.
- We never quote in-sample-era hit rates, only the latest genuinely out-of-sample window, with its base rate.
- We never quote anything from the sub-$5M market-cap universe; those prints aren't investable.
- We never present an early-labeled source as evidence; recompete and sentiment graduate when their outcomes accrue, not before.