somnial
- Karma
- 14
- Created
- 10 months ago
Recent Submissions
- 1. ▲ What happens when a GPU reads memory? (blog.doubleword.ai)
- 2. ▲ The case for disaggregated LLM serving (blog.doubleword.ai)
- 3. ▲ On-the-fly snapshot compression for elastic inference at scale (blog.doubleword.ai)
- 4. ▲ NVLink, NVSwitch, and All That (blog.doubleword.ai)
- 5. ▲ The Anatomy of an Instruction Pipeline Hazard (hiraditya.github.io)
- 6. ▲ Width vs. Depth: Speculating on the Margin (blog.doubleword.ai)
- 7. ▲ Pushing memory bound CUDA kernels past the speed of light with data compression (fergusfinn.com)
- 8. ▲ Speculative KV coding: ~4× losslessly compressed KV cache using a small model (fergusfinn.com)
- 9. ▲ 70x faster cold(ish) starts for SGLang (fergusfinn.com)
- 10. ▲ LLM powered data structures: A lock-free binary search tree (fergusfinn.com)