Methodology

How we audit

What Rigor tests, with which threshold and which sources, what each label means and what it does not do. The figures on this page are the ones the engine uses.

Independence

  • Rigor sells no robots, signals, courses or funded accounts, and its pages carry no affiliate links.
  • The report costs the same whatever class it gets: a better class never costs more.
  • We do not trade, hold money or keys, or connect to any broker.

Six questions, one threshold each

Is the result distinguishable from chance?

Probabilistic Sharpe ratio (length, skew and kurtosis) and a stationary block bootstrap of the returns.

What it takes to pass

PSR ≥ 0.95 and the bootstrap's 5th percentile Sharpe above zero; weak from PSR 0.80.

Does it survive the number of trials?

Deflated Sharpe ratio at the largest of the declared trials, the uploaded variants and the MT5 optimisation passes; PBO by combinatorial cross-validation when variants are uploaded.

What it takes to pass

DSR ≥ 0.95 and PBO below 0.50; weak from DSR 0.50.

Does it survive trading costs?

Every trade re-costed at 1x, 2x and 3x the cost, and the break-even cost.

What it takes to pass

Still positive at 3x the reference cost.

Does the out-of-sample stretch hold?

Sharpe after the out-of-sample start you declare, and its gap to the in-sample Sharpe.

What it takes to pass

Out-of-sample Sharpe ≥ 0.5 and a gap of at most 1.0.

Is the data sound?

35 red flags: duplicates, spikes, frozen marks, martingale, grid, deposits, backtest modelling and more.

What it takes to pass

No red flag. A serious flag fails the dimension; a warning makes it weak.

Does it beat what you could have held instead?

Excess return, drawdown ratio and information ratio against the benchmark you upload.

What it takes to pass

Positive excess return and information ratio, with a drawdown at most 1x the benchmark's.

How the A to D class is set

  1. A

    Statistics and number of trials pass; costs, out-of-sample and benchmark pass or do not apply; the data has no serious or warning flags.

  2. B

    Statistics pass, the number of trials passes or was not declared and nothing fails, but costs, out-of-sample, benchmark, data quality or the number of trials still need measuring or strengthening.

  3. C

    One dimension fails, or statistics or number of trials are weak.

  4. D

    The data or the statistics fail, or two dimensions or more fail.

What each label means

  • MeasuredWe computed it from your file.
  • DeclaredYou or your platform stated it; we cannot check it.
  • Not measuredData to measure it was missing, and the report says which.

The red flags we check

  • Too few observations
  • Zero or negative equity
  • Duplicated timestamps
  • Timestamps out of order
  • Unreadable rows
  • Returns without variation
  • Frozen marks
  • Extreme jumps
  • Implausible Sharpe ratio
  • Large gaps between rows
  • Zero declared costs
  • Fewer trials declared than the files show
  • Unreadable trades
  • Reported per-trade result does not match
  • Size grows after losses (martingale)
  • Grid or averaging down
  • Many positions open at once
  • Hidden floating drawdown
  • Many small wins and large losses
  • No sign of a stop loss
  • Result carried by a few trades
  • Trades outside the curve's dates
  • Trades that do not move with the curve
  • Printed final balance does not match the trade ledger
  • Curve and trades have an unexplained difference
  • Heuristic balance-chain signal
  • The percentage gain does not reflect the money
  • Deposits in a deep drawdown
  • Open loss the balance does not show
  • Backtest run on a coarse price model
  • Incomplete price history in the test
  • Settings on a lone peak
  • The optimisation does not hold in the forward period
  • The result fades in the recent period
  • The report header does not add up

Reproducible

  • Every file is identified in the report by its SHA-256 fingerprint.
  • Resampling uses a fixed seed the report prints: the same file with the same declarations gives the same numbers.
  • The report prints the engine version that produced it.

What it does not do

  • It does not predict future results: it measures the evidence in the data you upload.
  • It reads the files as they arrive; it does not check them with the broker.
  • It does not recommend buying, selling, copying or investing in anything.
  • A report is not legal, tax or investment advice.

Sources

  1. Bailey, D. H. y López de Prado, M. (2012). The Sharpe Ratio Efficient Frontier. Journal of Risk 15(2).
  2. Bailey, D. H. y López de Prado, M. (2014). The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting and Non-Normality. Journal of Portfolio Management 40(5).
  3. Bailey, D. H., Borwein, J., López de Prado, M. y Zhu, Q. J. (2017). The Probability of Backtest Overfitting. Journal of Computational Finance 20(4).
  4. Politis, D. N. y Romano, J. P. (1994). The Stationary Bootstrap. Journal of the American Statistical Association 89(428).

Public data we use

  • The Federal Reserve's exchange rates, the US Treasury bill and US consumer prices: FRED, Federal Reserve Bank of St. Louis. The VIX is Cboe Global Markets', through FRED.
  • The euro overnight rate (€STR, through FRED) and, before October 2019, the ECB's main refinancing operations rate (the fixed rate or, in the variable-rate tenders, the minimum bid rate) until October 2008 and the ECB's deposit facility rate after. Source: ECB statistics; this data is available free of charge on the ECB's website (ecb.europa.eu).
  • The sterling overnight rate (through FRED): SONIA data licensed under the Open Government Licence v3.0 and copyright the Governor and Company of the Bank of England.
  • Canada's overnight rate (CORRA): Bank of Canada; we convert it to an annual yield, and this data is available free of charge at bankofcanada.ca.
  • Brazil's monthly Selic rate: Banco Central do Brasil, series 4189, under the Open Database License (ODbL).
  • Policy rates of Mexico, Japan, Switzerland, Australia, New Zealand, India, South Africa, South Korea, Sweden, Norway, Denmark, Poland, Czechia, Hungary, Romania, Iceland, Türkiye, Israel, Saudi Arabia, Indonesia, Thailand, Malaysia, Chile, Colombia and Peru. Source: BIS (Bank for International Settlements). They are each central bank's official rate, not market rates.
  • Consumer prices for the euro area and Switzerland: Eurostat.
  • Consumer prices for the United Kingdom: Office for National Statistics, licensed under the Open Government Licence v3.0.
  • Consumer prices for Canada: Bank of Canada (Statistics Canada's CPI); this data is available free of charge at bankofcanada.ca.
  • Consumer prices for Brazil: Banco Central do Brasil (IBGE's IPCA).
  • Consumer prices for Mexico. Source: INEGI, Índice Nacional de Precios al Consumidor (INPC); we use it to take inflation out of the balances.
  • Consumer prices for Japan: created by editing the Consumer Price Index (Statistics Bureau, Ministry of Internal Affairs and Communications), through e-Stat.
  • All are read when the report is made and none changes the class.
  • We show no S&P 500, Nasdaq 100 or bitcoin closes: no public source allows their reuse in a paid report. The historical falls in the crisis table are fixed facts, not data we read.