JournalEdge Computing

Edge Inference Is a Memory Bandwidth Problem

Teams optimise FLOPs and wonder why the model is still slow on device. For single-batch generation, arithmetic was never the bottleneck.

25 Mar 20268 min readQuantization · Latency · On-Device

The instinct when a model runs slowly on constrained hardware is to reduce its arithmetic. Prune layers, shrink hidden dimensions, count FLOPs. Then the optimised model ships and is barely faster, and nobody can explain it.

Every token reads the whole model

In single-batch autoregressive generation, producing one token requires streaming essentially every weight from memory into the compute units. The arithmetic per weight is trivial — a multiply and an add. The transfer is not.

tokens/sec  ≈  memory bandwidth ÷ bytes per forward pass

7B model, fp16   → ~14 GB per pass
device bandwidth → ~50 GB/s
ceiling          → ~3.5 tokens/sec

same model, 4-bit → ~3.5 GB per pass
ceiling           → ~14 tokens/sec
A rough ceiling you can compute before writing any code

Quantisation gave a 4× speed-up without removing a single operation. It moved less data. That is the whole mechanism, and it is why weight-only quantisation is so effective on device while aggressive pruning often disappoints.

What follows from that

  • Measure achieved bandwidth, not utilisation. A device at 30% compute utilisation and 95% bandwidth is saturated.
  • Prefer weight-only quantisation before structural surgery — it targets the actual constraint and preserves accuracy better.
  • Watch the KV cache. It grows with context and competes with weights for the same bandwidth; long contexts degrade throughput even when the model fits.
  • Batching helps servers and does nothing for a single user on a phone. On-device economics are genuinely different.

On the edge, the question is rarely 'can this device do the maths'. It is 'how fast can this device read its own memory'.