4/ So we amortize the synthesis.
We train one small per-layer module (a Perceiver) once, against a frozen base model. At inference, learned latent queries cross-attend the full cache and write out compact keys & values. There are no reference queries, no gradient steps, no