JournalEdge Computing
Edge Inference Is a Memory Bandwidth Problem
Teams optimise FLOPs and wonder why the model is still slow on device. For single-batch generation, arithmetic was never the bottleneck.
The instinct when a model runs slowly on constrained hardware is to reduce its arithmetic. Prune layers, shrink hidden dimensions, count FLOPs. Then the optimised model ships and is barely faster, and nobody can explain it.
Every token reads the whole model
In single-batch autoregressive generation, producing one token requires streaming essentially every weight from memory into the compute units. The arithmetic per weight is trivial — a multiply and an add. The transfer is not.
tokens/sec ≈ memory bandwidth ÷ bytes per forward pass
7B model, fp16 → ~14 GB per pass
device bandwidth → ~50 GB/s
ceiling → ~3.5 tokens/sec
same model, 4-bit → ~3.5 GB per pass
ceiling → ~14 tokens/secQuantisation gave a 4× speed-up without removing a single operation. It moved less data. That is the whole mechanism, and it is why weight-only quantisation is so effective on device while aggressive pruning often disappoints.
What follows from that
- Measure achieved bandwidth, not utilisation. A device at 30% compute utilisation and 95% bandwidth is saturated.
- Prefer weight-only quantisation before structural surgery — it targets the actual constraint and preserves accuracy better.
- Watch the KV cache. It grows with context and competes with weights for the same bandwidth; long contexts degrade throughput even when the model fits.
- Batching helps servers and does nothing for a single user on a phone. On-device economics are genuinely different.
On the edge, the question is rarely 'can this device do the maths'. It is 'how fast can this device read its own memory'.