Volume quality methodology
Methodology v2 · last methodology review 2026-07-16 · facts verified 2026-07-16 · dataset collection started 2026-07-16.
PerpFinder does not claim that a low score proves wash trading, manipulation or fraud. The model measures how strongly reported activity is supported by selected observable market-quality signals. High turnover, market-maker programs and off-book flow can all legitimately produce high reported-to-support ratios. A weak relationship means additional verification is required — nothing more.
What is measured (and what is not)
- Measured: venue-reported 24h perpetual volume (venue-native tickers — not aggregator estimates), reported flagship BTC/ETH market volume, order-book median spread, buy-side VWAP slippage at standardized sizes, largest standardized order that fills, market-count and top-market concentration where the per-symbol breakdown allows it.
- Not measured in v1: depth in USD at fixed bps windows (the engine measures fill capacity at sizes instead — a depth-window layer is planned), sell-side slippage (buy-side walk only; sell requires the engine's bid-walk which is being rolled out), trade counts, temporal volume patterns (requires a mature series), and spot markets (separate future dataset — never blended into these rows).
- No adjusted dollar figure: v1 deliberately publishes no “normalized volume” dollar estimate. A multiplicative “support factor” without a validated model would be a fabricated number with fake precision; signals ship raw until a defensible model exists, and any future estimate will be labeled a PerpFinder estimate with its formula published here first.
Market selection and quote normalization
- Universe: CEX venues covered by BOTH our venue-native reported-volume feed and the order-book cost engine in the same sweep. Venues visible in only one source appear with reduced data confidence, not a worse verdict.
- Markets: each venue's flagship linear BTC and ETH perpetual (USD/USDT/USDC-quoted treated as one USD-class quote; the engine already targets each venue's main linear contract). Inverse contracts are out of scope in v1. We never compare an illiquid alt market on one venue against a flagship market on another.
- Stablecoin quotes are not FX-adjusted; a hard depeg beyond ±2% would make USD-class equivalence unsafe and pause affected comparisons (the engine's cross-venue median-price invariant catches this).
Signals and formulas
- Median spread: median of the venue's BTC and ETH flagship-market spreads (bps) in the sweep.
- Standardized slippage: the engine's VWAP-vs-best-ask walk at $10k / $100k / $500k / $1m (buy-side). Levels beyond 2% of best ask are junk-filtered; fills below 99.5% flag insufficient.
- Max standardized fill: the largest scenario size that fills on BOTH flagship markets. Both sides of the pair must support it — one deep market cannot carry a venue.
- Reported-to-support ratio: reported BTC+ETH main-market volume ÷ max standardized fill. Dimensionless and comparable across venues because the procedure is identical. Percentile ranks are winsorized at p5/p95 so a single extreme venue cannot own the scale. A high ratio is an observation, not an accusation.
- Data confidence (0–100): +30 reported volume available · +20 per-market breakdown available · +30 both flagship books observable · +20 full scenario ladder observable. Availability of OUR observation, kept strictly separate from any market-quality judgement. Missing data lowers confidence, never “quality”.
- No composite market-quality score exists in v1. If one is ever added it will publish its weights, inputs, sensitivity analysis and methodologyVersion here before shipping.
Collection, storage and missing data
- Sweeps run every 6 hours. Each sweep stores the verbatim inputs (reported-volume payload + full slippage ladder) as an append-only compressed envelope with
sourceTimestampandingestedAtkept separate — every derived row is reproducible from raw. - Writes are idempotent (bucket-keyed) with quality-priority upsert: a worse retry never replaces a better stored sweep; different methodology versions coexist.
- A venue failing in a sweep is isolated — other venues' rows are unaffected; the sweep is marked partial when coverage drops.
- Missing values render as “—/Unavailable”, never zero. Ratios with missing or zero denominators are not computed.
- Trend and temporal-pattern analysis is withheld until the series is mature (collecting → preliminary → established, requiring both age and observation count). History starts at our first stored sweep — no synthetic backfill.
- Storage growth: ~0.5 MB compressed per sweep ≈ 2 MB/day; aggregates and raw sweeps are both kept; no job deletes measurements.
Unverified volume
A venue reports its own 24h volume. PerpFinder normally checks that figure against two things we can observe ourselves: the venue's open interest, and the size its order book can absorb. When neither check is possible, the reported figure has nothing to stand on, and it does not enter a PerpFinder aggregate.
The rule. A venue is held out as unverified when it is listed in the unverified tier (lib/volume-verification.ts), or when it is a CEX with no readable open-interest feed whose reported 24h volume is above ten times the median volume-to-open-interest ratio of the venues that publish both. The list is the arm in force today; the ratio arm needs an open-interest figure the daily venue snapshot does not yet carry.
What the reader sees.The reported figure stays on the venue page and in every table, muted and labelled “unverified”. It is held out of the headline volume total, of the fee total, of every market share, of every rank and of the volume ranking chart. Each entry carries its own dated observations — the missing feed, the measured order-book depth against the reported figure, and the venue's own volume history.
Why it is separate from the spike rule. The spike rule holds a venue out while its 24h print sits above three times its own 7-day median, and returns it when the print settles. That is right for a one-day print and wrong here: once a large figure becomes the venue's own median, the spike clears and a number we still cannot check walks back into the total.
An entry is a measurement, not a verdict. PerpFinder does not claim that an unverified figure is wrong, and makes no claim about why a venue's reported figure and its observable market differ. An entry stays until an editorial review removes it, which needs a feed or a book we can read.
Limitations
- Order books are sampled snapshots; liquidity between sweeps is unobserved (persistence statistics arrive as the series matures).
- Visible depth is not guaranteed executable depth, and genuine internalized or off-book flow inflates reported-to-support ratios without any misreporting.
- Reported per-symbol breakdowns cover each venue's top markets only; concentration figures state their coverage.
- Buy-side slippage only in v1; bid/ask asymmetry is planned.
Corrections and attribution
Errors are handled under the correction policy. Pipeline: PerpFinder Data Team · editorial: PerpFinder Research · review: PerpFinder Editorial Review. Sources: venue public ticker and order-book APIs, measured by PerpFinder's own engine (methodology on /how-we-test and /data-definitions).
PerpFinder Research
Editorial TeamEditorial team tracking 100+ perpetual futures venues with live on-chain and exchange data.