RIS
github.com
2 threads
RIS reduces self-attention complexity to $O(N \log N)$ using sparse stochastic geometry that fits within commodity memory limits
RIS-Kernel: A Model-Agnostic Architecture for Long-Context LLM Inference via Sparse Attention