Everyone else sells you the backtester. This is the referee: a Sharpe deflated against a trial count the server keeps and you cannot edit, costs pinned at the exchange's real fee, and a holdout that opens once, ever. Ninety-four tools whose job is to take your best result away from you.
You still write the strategy and you still run the backtest. The only thing you lose is the ability to forget the ones that failed.
{ "mcpServers": { "quantbase": {
"url": "https://mcp.quantbase.live/mcp",
"headers": { "Authorization": "Bearer <your key>" } } } }
A strategy you are tuning — in Python, in a notebook, in an agent.
A Sharpe deflated against every variant you tried — counted by the server, not by you. Plus PBO, a confidence interval, and a holdout that opens once.
Needs Claude, Cursor or any MCP client.
A backtest someone else is showing you — a signal group, a bot seller, a pitch.
Their Sharpe, deflated against how many things they tried. Type two numbers and watch what is left of it.
Check one now → No account.
An idea and no data. You want to know if it survives before you fund it.
The server brings the prices and the exchange fees. You bring the idea. 200 trials a month, free.
No card. Email only.
Not for you if you are picking coins to buy, or want signals. This server emits none — it only tells you whether a result you already have is real.
Your client talks to one server. The server owns the data, the splits, the statistics and the trial count, and reaches out to the exchange for prices and fees.
Thirty equity curves, every trade in them random. The best one is simply the top of the fan, and it is the one you would have funded.
Any test you re-run until it passes will pass. That includes significance tests and Monte Carlo reshuffles: they measure the winner and cannot see the losers you threw away.
A Sharpe of 1.5 sounds like a grade. It is a measurement, and it means nothing until you know how many times you measured.
A Sharpe ratio is only good relative to the best of N tries. Try thirty variations, keep the winner, and you have cleared a much lower bar than you think. N is the whole correction, so the server keeps its own append-only count. You cannot pass it in, lower it, or delete a trial.
The best of N attempts with no edge is expected to score about σ·√(2 ln N). Everything the server computes starts from this line: the deflated Sharpe is the chance your result beat it.
N is the entire correction, which is exactly why it is the one number you are not allowed to type in.
A moving-average family on BTC returned a Sharpe of 0.95 in-sample, the best of 28 recorded training runs. Run unchanged across ten assets, the pooled figure is negative. The gap between those two numbers is the product.
Breadth is the cheapest honesty test there is. A rule with a real edge survives being moved; a rule fitted to one asset's history does not, and no amount of in-sample Sharpe tells you which you have until you move it.
Ten assets on one exchange are not ten tests. The next section is about how few they really are.
Crypto moves together. Ordered by cluster, the correlation matrix has three bright blocks and little else, so twenty-seven assets buy you three independent looks, not twenty-seven.
Costs are checked before a run, not after. Most strategies die here, and it is cheaper to find out now.
The matrix is ordered by cluster and drawn to illustrate the structure; the 0.52, 57% and 3.0 are measured on the server's own data. The fee is Coinbase's published taker rate at the entry tier, which is what a new account actually pays.
Buy the fee schedule before you buy the idea. The round trip decides how much edge a strategy can ever keep.
It arrives with the trial count that produced it, the deflated figure, and the confidence interval, drawn at the same weight when that interval straddles zero as when it does not.
The server will not hide an interval that crosses zero, and it will not let a strategy be re-tested until the trial count says so. There is no route from a failed gate to the holdout.
The holdout is the moment the backtest stops being yours and becomes evidence.
A fixed span of data supports a maximum number of honest tries. Past it, no result can be told apart from the best of many coincidences. These are this server's own families, right now — including the one that has already spent more than its data can cover.
Free needs no card. Holdouts are metered separately: each one opens exactly once per strategy, ever.
Enough to take one honest idea through the whole gauntlet.
Start freeFor running a family properly: sweeps, splits, ten-asset pooling.
Choose ProShared ledger across a desk, with per-tenant isolation.
Choose TeamOne backtest, one sweep point, or one cross-section run. Every trial permanently raises the bar your own results must clear, so it is the scarce thing worth pricing, and a smaller sweep buys you a lower hurdle as well as a smaller bill.
The referee says no far more often than yes. That is the product, not a flaw in it.