Mostik.ai – latent communication between AI models
mostik.ai
1 thread
Conjecture: Does GLM only handle the prefill stage, compress its output hidden vectors (trained?), and then send them to Qwen for decoding?
Conjecture: Does GLM only handle the prefill stage, compress its output hidden vectors (trained?), and then send them to Qwen for decoding?