SwiGLU has no exploitable contextual sparsity
Contextual sparsity — predicting which FFN neurons will be active and skipping the rest — is a well-established acceleration technique. It does not transfer to Llama-3.
Concentration
Channels required to hold 90% of the activation magnitude, out of 8192:
| layer | channels | % |
|---|---|---|
| 0 | 3772 | 46.0 |
| 7 | 4493 | 54.9 |
| 14 | 4470 | 54.6 |
| 27 | 3558 | 43.4 |
Half. There is no small subset to skip.
Stability
Overlap of the top-k active set between consecutive tokens:
| layer | k=256 | k=512 | k=1024 |
|---|---|---|---|
| 0 | 10% | 13% | 18% |
| 14 | 29% | 27% | 31% |
| 27 | 38% | 43% | 56% |
The set churns 70-90% every token, so it cannot be predicted cheaply from the previous step either.
Why
sat_gate(g)·u is smooth. Every one of the 8192 channels comes out nonzero, just small. ReLU zeroes 90%+ of channels outright — which is where the published results come from. Llama-3 traded that sparsity away for quality.
Best case with a perfect oracle is 2x on down_proj, 25% of a layer, so about 12% of total weight traffic — and the oracle cannot exist at 20% overlap.
The route that would work is ReLUfication: fine-tune the activation back to ReLU and recover the sparsity. That is a training problem, not an inference one.
← Back to the library