Skip to content
SimicX
Quant Research · Alpha Evaluation Guide

How to tell whether an alpha actually works

An alpha signal is easy to generate and easy to fool yourself with. This guide walks through how a trading alpha is evaluated in practice: the two families of alpha, the scorecard each requires, and the checks — realistic costs, out-of-sample discipline, correction for multiple testing — that separate a real edge from a lucky backtest. Nothing here is a capital recommendation; it is a framework for reading any alpha claim, including ours, with the right degree of skepticism.

0
Alpha families
0
Evidence tiers
0
Defined metrics
0
Interactive labs

Cross-sectional alpha

Ranks a broad universe at each date; the bet is relative — which names beat which. Judged on cross-sectional rank IC, quantile spreads and the P&L of a simple long/short book, with turnover, cost and multiple-testing checks on top.

Time-series forecast alpha

Forecasts each asset's own future return, traded as a concentrated long/short book of a few names. Because the forecast maps directly to the position, it is judged on the out-of-sample P&L of that book, net of costs and deflated for multiple testing.

◆ The core principle

Judge each strategy on what it actually trades, then subtract what it would really cost to run. A cross-sectional alpha makes money by ordering names correctly on each date; a time-series alpha makes money by calling each asset's own direction over time. Borrow one family's measure for the other and it misleads: a ranking score across a handful of names is mostly noise, and a strong ranking score says nothing about whether a directional book makes money. To compare two candidates fairly, put both through the test that suits them — after costs, after allowing for how many ideas were tried before this one, and on data neither has seen.

01 · Foundations

Two families of alpha, two scorecards

An alpha begins as a single score per asset — per bar for an intraday strategy, per rebalance date for a multi-day one. How that score is turned into a verdict depends on what the signal is designed to capture, and the process differs between the two families below. Part A covers the ranker's scorecard, Part B the forecaster's.

CROSS-SECTIONAL · across names, one daterank the universe → IC = corr across namesTIME-SERIES · one asset, over timeforecast vs realized over time → corr over time

Left: the ranker's skill is whether high scores out-ranked low scores at each date. Right: the timer's skill is whether the forecast called each asset's own path. The two skills answer different questions, so they get different scorecards.

◆ Why the distinction matters

The Fundamental Law of active management, , states the problem precisely. For a ranker, breadth is the number of names and IC is the cross-sectional rank correlation. For a timer, breadth is the number of independent time periods and the relevant correlation runs over time — and effective breadth shrinks under serial and cross-asset correlation, . Apply the wrong definition and the headline number is meaningless.

Sign in to continue

See the full evaluation framework

You've seen the two families and why each needs its own scorecard. The rest of the guide — the eight questions, the metric deep-dives, and the interactive labs — unlocks free the moment you sign in.

  • The eight questions every full evaluation answers
  • Part A — the cross-sectional scorecard: rank IC, quantile spreads and typical healthy ranges
  • Part B — the three-tier time-series verdict and forecast-quality diagnostics
  • Parts C & D — costs, capacity and the frequency rulebook
  • The capital decision, six interactive labs, and the full glossary

Checking your access…