5.1 Market Making Is Not Just Collecting the Spread
Posting a bid and ask is the easy part. The money comes from forecasting fair value, dodging informed flow, and skewing thousands of weak alphas into passive quotes at zero cost.
Pillar
Short-horizon, execution-aware edge. Fair value, markouts, OBI, microprice, spreads and skew, adverse selection, fill probability, and maker/taker economics.
Posting a bid and ask is the easy part. The money comes from forecasting fair value, dodging informed flow, and skewing thousands of weak alphas into passive quotes at zero cost.
A market maker's quote is three stacked decisions: fair price, spread, skew. Fair price carries more than the other two combined, and skew, the academic favorite, is the trivial one.
Fair value is the price that gives good markouts when you quote around it. If your losses realize 750 ms after a fill, fair value is where price will be then, not today's mid.
The markout asks one question of every fill: what did the price do next? It ignores the spread you collected and the cleverness you intended, exposing toxic flow as a curve that sinks.
A maker's forecast is graded only on the fills it gets, and counterparties hand you the adverse ones. A 40% model across all moves can be a 30% model on the trades you actually make.
Quote both sides with no view and informed takers hit whichever side is about to move, leaving you the losing inventory. That is adverse selection, the permanent maker-versus-taker battle.
Holding some inventory is benign. Being loaded with toxic inventory by someone who knows it'll move against you kills the book. Skew exists to stop the loading, not to tidy variance.
Skewing is the academic favorite and the desk's afterthought. Make one side wider so it fills less, lean harder against aggression, and skip the stochastic control. Keep it simple.
Blow your spreads out when volatility does, because a fixed markup that was safe in a calm tape becomes a gift to informed traders in a moving one. Around scheduled news, stop quoting entirely.
Order placement, the exact tick you rest on, matters more than skewing. Sit just in front of a large sturdy order so takers can't push the price through your fill, and your markouts improve.
The book is a landscape of walls and gaps. Average density over minutes and place in front of shelves that stay stocked, weighing the edge of sitting behind size against the aggression to reach it.
Half the big orders in crypto are spoofs that vanish on approach. Lean on one and the price runs through your fill. Filter for sturdy size, and trade the turn a spoof's disappearance creates.
Order book imbalance, bid size minus ask size over their sum, is the strongest simple microstructure feature. Price rolls toward the thin side, and the signal lives in the tails, so fit a spline.
The plain mid ignores resting sizes. The weighted mid and microprice lean the center toward the thin side, where price is headed, giving a maker a fair value that already respects the book.
Trade flow reads liquidity being taken, not posted. Signed and summed, it predicts price because trades are autocorrelated: buys follow buys. And unlike book imbalance, executions can't be spoofed.
A quote's PnL is easy; its fill probability is the hard part. Build a CDF of taker order sizes and read off the chance the next trade is big enough to clear the queue ahead of you plus your own size.
TWAP and VWAP are not indicators to cross over. They are execution benchmarks and the slicing algorithms built to hit them, the machinery for moving a large order without blowing out the price.
A half-bp signal is garbage to a taker who pays the spread and money to a maker who collects it. That cost flip turns thousands of weak alphas into the bulk of a maker's PnL.
One signal, two businesses. The taker pays the spread and needs a strong edge; the maker earns it and thrives on weak ones across a wider universe at higher frequency. The spread divides them.
A maker shouldn't quote everything. Track the EWMA markout of every trade in a symbol, quote only the non-toxic names, and scale size by markout. Selection alone can turn a flat system profitable.
On a follower crypto venue, your local mid is stale. Build fair value from the leader: regress the basis against Binance's mid, blend global and local prices, and respect the stablecoin rates.
MSE on variance is the wrong ruler for vol forecasts. QLike is optimal without knowing the distribution and punishes underestimating vol more than overestimating, matching the maker's adverse-selection cost.
Bar research assumes the clock drives the market. It doesn't, events do. Fixed rebalance times are front-runnable and bars hide execution seasonality. Screen on bars, validate on event-driven sims, trade on events.
The cheapest way for an MFT shop to enter a small position is to act like a market maker: feed your quoting engine a phantom short so it skews and fills you long, collecting the spread instead of paying it.
The Fourier transform fails on price but nails a robot, because a wash bot or metronomic TWAP is the clean periodic signal price never is. Bin trades, FFT the counts, and a sharp spike is automation. The catch: your own equal-interval TWAP makes that spike, so randomize it.
Volatility predicts the size of a move, not its direction, so it can't trade alone. Multiply it by momentum and it works, until extreme vol gains its own sign, down for equities and up for gold.
Got forecasts on different clocks? Don't pick the biggest. Normalize each to edge-per-second on its markout curve (30 bps/60s = 0.5 vs 2 bps/1s = 2) and average the curves to blend fast and slow edge into one.
A quote stamped newer than light can travel isn't fast, your clocks disagree. Fix chrony first, then use an empirical latency floor (half the fastest ping) to catch and correct cross-venue drift.
You can't simulate a limit fill that never happened. Maker/taker on historical trades works only on thin venues; pure making is just prod tuning. And Avellaneda-Stoikov is risk times root holding time.
Crypto vol runs on a 24-hour clock, spiking at 14:00 and 00:00 UTC from session opens and midnight rebalancing. GARCH and HAR have no clock term, so they quote too tight into the predictable bursts.
SAR fixes vol seasonality by adding lags at 24 and 168 hours to an AR model, fit by OLS. The daily lag injects the midnight and 14:00 bumps a plain AR(1) misses, so you stop quoting tight into them.
Slice average return and risk by UTC hour and a time-of-day premium appears: long BTC at midnight, short into the early hours. Then four fills of fees eat the 14 bps gross, which is why you treat it as a tilt, not a strategy.
The exchange 24h change rolls, so the headline jumps when an old crash drops out of the window, not when price moves. Uninformed traders chase the mirage; front-run them by reading the hour about to exit.
Volume does not tell you direction, it tells you whether the move sticks. High volume means continuation, low volume means reversion. Use it as the regime switch on a directional signal, not as the signal.
When New York opens, a volume surge hits crypto. From 13:30 to 15:00 UTC, if volume keeps rising, ride the sign of the first half hour to the close of the window. Track the real open, not a frozen timestamp.
Imbalance tells you how much size rests in the book. Arrival, cancellation, and update rates tell you how fast it churns, and the cancellation rate, measured right with the trade feed, is your spoofing alarm.
A maker repricing every few hundred milliseconds needs tick-scale volatility, not one-minute bars. Measure it three ways: std of traded prices, the book's churn rate, or the volatility of your own fair price.
Market orders are autocorrelated, roughly AR(1): buys follow buys in clustered spikes. Don't post offers into a buy run. And beware, OLS underestimates phi, so your model clears you to requote half a run too early.
Crypto books break the textbook: negative spreads are data artifacts, blowouts are wipeouts from large orders clearing levels, and more resting size means a tighter spread because deep books are competitive books.
Your spread must track volatility, but how? Bucket volatility into three quantile regimes with fixed widths, or interpolate between them (linear, spline, sigmoid) so the spread glides instead of jumping at the boundary.
A big maker's PnL isn't the spread, it's positional: skewing thousands of weak alphas into passive quotes. Zero cost makes half-bp signals tradeable, and ensembling them cancels noise into the bulk of the profit.
Market impact is the cost you simulate, not look up: walk the book level by level, slippage = (avg - mid)/mid x 1e4, then fit a·x^b. The exponent near 0.5 is the square-root law, and it lets you price sizes past the visible book.
The exchange never hands you the order book, just a snapshot plus a firehose of deltas. Fold them in order using a hashmap for O(1) edits and a sorted tree for best bid/ask. Miss one sequence number and every feature silently lies.
The naive lead-lag trade enters B after A moves and exits on a timer. That wastes the edge. Use the slope as a forecast of B's move, exit the instant B hits it, and hold to the horizon only as a backstop.
Lead-lag edge lives in the big moves, so quantile-regress on the tail, not the mean. Then label the cause from trade size and the liquidation feed: impact follows clean, liquidations snap back, and scheduled news is a cue to widen quotes, not to trade.
Hyperliquid runs fully on-chain, so a node captures every order event: wallet IDs, counterparty inventory, and the ~89% of orders that are rejected and invisible in LOBSTER-style data.