Volume Quality Methodology
Methodology v1 · last methodology review 2026-07-16 · facts verified 2026-07-16 · dataset collection started 2026-07-16.
PerpFinder does not claim that a low score proves wash trading, manipulation or fraud. The model measures how strongly reported activity is supported by selected observable market-quality signals. High turnover, market-maker programs and off-book flow can all legitimately produce high reported-to-support ratios. A weak relationship means additional verification is required — nothing more.
What is measured (and what is not)
- Measured: venue-reported 24h perpetual volume (venue-native tickers — not aggregator estimates), reported flagship BTC/ETH market volume, order-book median spread, buy-side VWAP slippage at standardized sizes, largest standardized order that fills, market-count and top-market concentration where the per-symbol breakdown allows it.
- Not measured in v1: depth in USD at fixed bps windows (the engine measures fill capacity at sizes instead — a depth-window layer is planned), sell-side slippage (buy-side walk only; sell requires the engine's bid-walk which is being rolled out), trade counts, temporal volume patterns (requires a mature series), and spot markets (separate future dataset — never blended into these rows).
- No adjusted dollar figure: v1 deliberately publishes no “normalized volume” dollar estimate. A multiplicative “support factor” without a validated model would be a fabricated number with fake precision; signals ship raw until a defensible model exists, and any future estimate will be labeled a PerpFinder estimate with its formula published here first.
Market selection and quote normalization
- Universe: CEX venues covered by BOTH our venue-native reported-volume feed and the order-book cost engine in the same sweep. Venues visible in only one source appear with reduced data confidence, not a worse verdict.
- Markets: each venue's flagship linear BTC and ETH perpetual (USD/USDT/USDC-quoted treated as one USD-class quote; the engine already targets each venue's main linear contract). Inverse contracts are out of scope in v1. We never compare an illiquid alt market on one venue against a flagship market on another.
- Stablecoin quotes are not FX-adjusted; a hard depeg beyond ±2% would make USD-class equivalence unsafe and pause affected comparisons (the engine's cross-venue median-price invariant catches this).
Signals and formulas
- Median spread: median of the venue's BTC and ETH flagship-market spreads (bps) in the sweep.
- Standardized slippage: the engine's VWAP-vs-best-ask walk at $10k / $100k / $500k / $1m (buy-side). Levels beyond 2% of best ask are junk-filtered; fills below 99.5% flag insufficient.
- Max standardized fill: the largest scenario size that fills on BOTH flagship markets. Both sides of the pair must support it — one deep market cannot carry a venue.
- Reported-to-support ratio: reported BTC+ETH main-market volume ÷ max standardized fill. Dimensionless and comparable across venues because the procedure is identical. Percentile ranks are winsorized at p5/p95 so a single extreme venue cannot own the scale. A high ratio is an observation, not an accusation.
- Data confidence (0–100): +30 reported volume available · +20 per-market breakdown available · +30 both flagship books observable · +20 full scenario ladder observable. Availability of OUR observation, kept strictly separate from any market-quality judgement. Missing data lowers confidence, never “quality”.
- No composite market-quality score exists in v1. If one is ever added it will publish its weights, inputs, sensitivity analysis and methodologyVersion here before shipping.
Collection, storage and missing data
- Sweeps run every 6 hours. Each sweep stores the verbatim inputs (reported-volume payload + full slippage ladder) as an append-only compressed envelope with
sourceTimestampandingestedAtkept separate — every derived row is reproducible from raw. - Writes are idempotent (bucket-keyed) with quality-priority upsert: a worse retry never replaces a better stored sweep; different methodology versions coexist.
- A venue failing in a sweep is isolated — other venues' rows are unaffected; the sweep is marked partial when coverage drops.
- Missing values render as “—/Unavailable”, never zero. Ratios with missing or zero denominators are not computed.
- Trend and temporal-pattern analysis is withheld until the series is mature (collecting → preliminary → established, requiring both age and observation count). History starts at our first stored sweep — no synthetic backfill.
- Storage growth: ~0.5 MB compressed per sweep ≈ 2 MB/day; aggregates are permanent, raw retained ≥90 days minimum.
Limitations
- Order books are sampled snapshots; liquidity between sweeps is unobserved (persistence statistics arrive as the series matures).
- Visible depth is not guaranteed executable depth, and genuine internalized or off-book flow inflates reported-to-support ratios without any misreporting.
- Reported per-symbol breakdowns cover each venue's top markets only; concentration figures state their coverage.
- Buy-side slippage only in v1; bid/ask asymmetry is planned.
Corrections and attribution
Errors are handled under the correction policy. Pipeline: PerpFinder Data Team · editorial: PerpFinder Research · review: PerpFinder Editorial Review. Sources: venue public ticker and order-book APIs, measured by PerpFinder's own engine (methodology on /how-we-test and /data-definitions).
PerpFinder Research
Editorial TeamEditorial team tracking 100+ perpetual futures venues with live on-chain and exchange data.