THE METHODOLOGY
The math. All of it.
A benchmark you cannot audit is a benchmark you should not trust. This page walks the entire pipeline — every formula, every guard, and every place the number is weakest — in the order the engine runs it. Nothing here is proprietary; the only private thing in the system is your data.
ENGINE v1.1.0 · THIS PAGE DESCRIBES THE SHIPPED ENGINE, NOT AN IDEAL
v1.1 moved open-option marks off intrinsic value onto a priced model, so some numbers you may have already seen changed when this page's own math changed under them — disclosed here, not smoothed over.
The question the number answers
Profit does not answer whether you can trade — a rising market pays everyone. The rating asks a narrower question: did your decisions beat the market's own risk-adjusted offer? EDGE is a signed handicap in percentage points per year — the annualized, risk-adjusted distance between what your account did and what blind exposure to the S&P 500 did with the same money on the same dates. Zero is the market. A positive number is edge the market did not hand out. A negative number is what blind exposure paid over the same window that your decisions did not.
Inputs: the broker's records, not your screenshots
The engine reads your brokerage's own activity record through a read-only connection: fills, deposits, withdrawals, dividends, fees, and option actions including assignment and exercise. From that stream it rebuilds your positions and cash day by day. Every activity carries the provider's identifier — or, where the broker omits one, a content-derived fingerprint — so a re-sync can never double-count a trade.
Before any rating is computed, the reconstructed account value is checked against the value the broker itself reports. If they disagree beyond tolerance, the engine refuses to rate rather than rate wrongly. The whole pipeline prefers a loud failure over a quiet lie.
Daily valuation
Each position is marked at its dividend- and split-adjusted end-of-day close, turning the account into a daily equity curve. The product displays derived analytics only — never price tables. Open option positions are priced with a Black-Scholes model (European exercise, r = 0) against the adjusted underlying close, using sigma from the underlying's realized volatility over the trailing 63 trading days, and the model price is floored at intrinsic value so it can never mark a position below what exercising it today would pay. Fewer than 20 prior closes and there isn't enough history to fit a volatility — the mark falls back to intrinsic, and the dossier discloses which of your options were priced which way. Realized options profit and loss stays exact throughout, from your fills, no matter how an open position was marked along the way. And when a required underlying price is missing, the engine refuses to value that day instead of interpolating silently.
Flow-adjusted time-weighted return
A deposit is not a return. Each day's return is computed with that day's external flows removed, then the daily returns compound geometrically — the standard defense against cash flows masquerading as skill.
The blind-SPY counterfactual
The benchmark is not an index quote — it is a full counterfactual account: your exact deposits and withdrawals, on your exact dates, executed into a dividend-adjusted S&P 500 proxy with zero decisions. Raw edge is your annualized TWR minus the counterfactual's. In the worked example used across this site (one consistent fictional trader): +15.7% against +11.1% — raw edge +4.6.
The risk penalty: downside-M²
Raw edge bought with reckless risk is not edge. The engine measures downside deviation — only losing days count — annualized over 252 trading days:
Your return is then rescaled to the benchmark's downside risk (a Sortino-style M², risk-free rate 0 by documented v1 convention), and the benchmark's annualized return is subtracted. In the worked example, downside deviation ran 13.9 against the counterfactual's 12.6 — already the larger of the two, so the floor changes nothing here:
The floor exists because a thin history understates risk — a lucky quiet stretch is not low risk. When the benchmark itself recorded no downside in the window, the ratio has no scale and the engine falls back to raw excess — the only defined comparison, and one that preserves the foundational invariant that an exact index clone grades 0. Leverage is deliberately not in this floor: its denominator would include cash, so a deposit would move the handicap. Leverage is punished where it actually shows up — realized drawdown — and reported in the decomposition; in the worked example, realized drawdown ran −12% against the counterfactual's −9%. Worked example: raw edge +4.6, risk penalty −1.5, EDGE +3.1 — tier OPERATOR.
THE EXPOSURE LINE
A separate line answers a separate question: how much of your return was just carrying the market? An OLS regression of your daily returns on the benchmark's paired daily returns (minimum 40 paired days) yields beta — how geared to the market your account behaved — and Jensen's alpha, the regression's own intercept, annualized arithmetically (×252, disclosed as such) at a risk-free rate stated as zero, the same v1 convention used above. R² reports how much of your day-to-day variance the market explains at all. None of this feeds EDGE's arithmetic — the exposure line is reported beside EDGE, never inside it.
In the worked example: beta 1.32, Jensen's alpha +1.9, R² 0.71 — exposure explains most of the account's day-to-day variance, and the alpha left over sits below raw excess return, which is what carrying the market rather than beating it looks like on this line.
Six tiers, fixed bands
Bands are fixed to the number, never to a population quota — a tier means the same thing at ten rated traders or ten thousand. The boundaries are calibrated against the Barber & Odean annualized quartiles (25th ≈ −8.8, median ≈ −1.7, 75th ≈ +6.0). Each band owns (lo, hi]: EDGE −8.0 is still EXIT LIQUIDITY; only strictly above +10 is THE HOUSE.






The reference percentile
Until 500 accounts hold verified ratings, your percentile is placed against published research rather than a population we do not yet have — principally Barber & Odean (2000), Table IV: net-of-cost, market-adjusted annualized returns of 62,439 discount-brokerage households. Placement is piecewise-linear interpolation between the published points; outside the observed range it clamps to 0.5 / 99.5 rather than extrapolate.
The switch itself is mechanical, not editorial. Every night the engine snapshots every account holding a published rating — cleared the same provisional gate every dossier clears, nothing more required — and counts it. The moment that count reaches 500, the next nightly snapshot becomes the reference table: the same piecewise-linear interpolation, the same 0.5 / 99.5 edge clamp, now built from rated traders instead of a 1990s academic sample. The switch will be announced as an upgrade, not slipped in.
Honest footnote: this note will look dated the day the switch happens — a disclosure written before an event always describes something that has already occurred by the time you read it after the fact. That is intentional. The flip is a recorded change, not a live status this page keeps current for you.
| PERCENTILE | RETURN vs MARKET |
|---|---|
| p1 | −58.3 |
| p5 | −29.4 |
| p10 | −19.2 |
| p25 | −8.8 |
| p50 | −1.7 |
| p75 | +6.0 |
| p90 | +16.8 |
| p95 | +25.8 |
| p99 | +53.3 |
Your reference percentile is built from published academic studies of retail trader performance, principally Barber & Odean (2000, The Journal of Finance): net-of-cost, market-adjusted returns of 62,439 US discount-brokerage households, 1991–1996, annualized from monthly figures. It is not a distribution of current PAYDOSSIER_ users. The underlying studies mix markets, eras, instruments, and units; where a figure is gross or not risk-adjusted, the methodology page says so. Zero-commission trading has changed cost structures since these samples were collected. Past cohort distributions do not predict any individual's results.
Where the points went: five habits
The decomposition attributes the gap between you and the counterfactual to five recorded habits. It is strictly backward-looking: it reports what each habit cost or earned in the window, past tense. It is not a coaching plan.
- SELECTION — what each closed round trip earned against blind SPY over the same holding days.
- EXIT-TIMING — the five weekdays after each exit: what the sold position recorded next.
- SIZING — whether the positions you sized largest recorded better or worse outcomes than the small ones.
- CONCENTRATION — portfolio concentration (HHI) above the five-equal-positions baseline of 0.2; each 0.1 above it cost one point in the v1 heuristic.
- LEVERAGE — exposure carried above equity, and what it amplified.
Honest footnote: the three trip habits are annualized rates while concentration and leverage are levels, and v1 ranks them on one scale. The ordering is a disclosed heuristic, not a theorem — treat the ranked list as "largest recorded drag first," not as precision.
When the engine refuses to grade
A rating stays provisional until the window holds at least 60 trading days and 10 closed round trips — below that, variance swamps skill. Accounts that recorded no positions, went dormant after a full withdrawal, or show only external flows get no grade at all. A broken broker connection flips the rating to a dated STALE badge; it never silently ages.
THE CONFIRMATION STAMP
Clearing the provisional gate earns a rating; it does not by itself earn the CONFIRMED stamp. The engine re-runs the same handicap pipeline across 1,000 replays of your own account and benchmark history, drawn by a stationary bootstrap — blocks averaging 5 trading days, so an autocorrelated stretch (a hot week, a drawdown) gets resampled as a stretch, not shuffled apart into days that never actually followed one another. Each replay lands in one of the six tiers, and the share landing in the tier the point estimate was awarded decides whether that tier is earned on this run.
A tier reached for the first time needs the majority outright: at least half the replays landing in the awarded tier, or the stamp stays PROVISIONAL — tier assigned, not yet defended under resampling. But once a tier has been CONFIRMED, that confirmation carries forward on every later recomputation for as long as the tier holds, even on a run where the replay share dips below half — one noisy week can't flap the stamp off. The carry breaks the moment the tier changes: the new tier starts PROVISIONAL and has to clear its own majority before it, too, can be marked CONFIRMED. The closed-trade gate above never stops applying — a confirmation cannot outlive an account that no longer meets it.
Honest line: accounts with thin or volatile histories stay provisional longer, sometimes a full quarter — the bootstrap's spread tracks your own variance, and a wide replay spread isn't a flaw in the confirmation, it's the confirmation correctly declining to certify what the data can't support yet.
Where the number is weakest
Open positions are priced by a model, not a live quote — Black-Scholes, European exercise, no volatility smile, risk-free rate held at zero. It can misprice anything the model's own assumptions don't hold for: American-style early exercise, skew, dividends timed inside the option's life. Fewer than 20 prior closes and it falls back to intrinsic value instead of guessing. Realized options profit and loss remains exact from fills regardless.
Idle cash genuinely dragged risk-adjusted returns in past windows — that is real economics, not a measurement leak, but it means a deposit left uninvested lowered subsequent grades.
A near-cash account in a falling market can grade absurdly well — it merely did not lose. The provisional gate absorbs most of this; the residual framing is disclosed here.
Brokers expose different history depth, and per-broker limits are still being verified against real accounts. Until that pilot closes, no claim is made about how much of your past connects instantly.
The reference distribution predates zero-commission trading (households, 1991–1996). Costs changed; behavior largely did not. The disclosure above carries this caveat wherever a percentile appears.
The bootstrap resamples the year you actually traded — blocks reordered, never invented. It cannot manufacture a regime your account never lived through: a crash your window missed stays missing from every one of the 1,000 replays. CONFIRMED means the tier held up under resampling of your own recorded history, not that it will hold in a year that hasn't happened yet.
Invariants: tested, not promised
These properties are pinned by the engine's test suite — a change that breaks any of them cannot ship:
SPY clone → EDGE exactly 0— replicating the benchmark earns no edge.all activities × k → identical rating— account size is invisible.deposit dated the rating day → Δ exactly 0— money arriving cannot move the number.deposit invested same day → Δ exactly 0— only decisions move it.SPY clone → EDGE exactly 0 · 100% TAPE READER tier odds— replicating the benchmark earns no edge, and zero variance across the replays means every one of them lands in the same tier; the confirmation gate is separate and still requires closed trades, so a clone that never traded stays provisional regardless.same input → identical rating, bootstrap included— the resampling PRNG is seeded from the account's own data, never the clock; recomputing the same history reproduces the same confirmation, replay for replay.
The sources
Every research figure in the product traces to one of these. Sources are listed with their weaknesses on purpose.
[1]
Barber, Brad M., and Terrance Odean (2000)
Trading Is Hazardous to Your Wealth: The Common Stock Investment Performance of Individual Investors
The Journal of Finance · Vol. 55, No. 2, pp. 773–806, Table IV p. 791
https://onlinelibrary.wiley.com/doi/abs/10.1111/0022-1082.00226[2]
Chague, Fernando, Rodrigo De-Losso, and Bruno Giovannetti (2020)
Day Trading for a Living?
SSRN Working Paper No. 3423101 · working paper, June 2020 revision
https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3423101[3]
Barber, Brad M., Yi-Tsung Lee, Yu-Jane Liu, and Terrance Odean (2009)
Just How Much Do Individual Investors Lose by Trading?
The Review of Financial Studies · Vol. 22, No. 2, pp. 609–632
https://academic.oup.com/rfs/article-abstract/22/2/609/1595677[4]
Barber, Brad M., Yi-Tsung Lee, Yu-Jane Liu, and Terrance Odean (2014)
The cross-section of speculator skill: Evidence from day trading
Journal of Financial Markets · Vol. 18, pp. 1–24
https://www.sciencedirect.com/science/article/abs/pii/S1386418113000190[5]
Bauer, Rob, Mathijs Cosemans, and Piet Eichholtz (2009)
Option trading and individual investor performance
Journal of Banking & Finance · Vol. 33, No. 4, pp. 731–746
https://www.sciencedirect.com/science/article/abs/pii/S0378426608002720[6]
Barber, Brad M., Xing Huang, Terrance Odean, and Christopher Schwarz (2022)
Attention-Induced Trading and Returns: Evidence from Robinhood Users
The Journal of Finance · Vol. 77, No. 6, pp. 3141–3190
https://onlinelibrary.wiley.com/doi/abs/10.1111/jofi.13183[7]
European Securities and Markets Authority (2018)
Product Intervention Analysis: Measure on Contracts for Differences
ESMA · ESMA50-162-215, 1 June 2018
https://www.esma.europa.eu/press-news/esma-news/esma-agrees-prohibit-binary-options-and-restrict-cfds-protect-retail-investorsWhat this is, and is not
The rating is backward-looking analytics computed on your own account's recorded history. It describes what happened; it predicts nothing and recommends nothing — no trades, no strategies, no securities. It is not investment advice. Dollar amounts are never shown publicly: the tier card carries the tier, the handicap's sign, the season, and the streak — never balances, never positions.