Settings

Theme

somnial

Karma
14
Created
10 months ago

Recent Submissions

  1. 1. What happens when a GPU reads memory? (blog.doubleword.ai)
  2. 2. The case for disaggregated LLM serving (blog.doubleword.ai)
  3. 3. On-the-fly snapshot compression for elastic inference at scale (blog.doubleword.ai)
  4. 4. NVLink, NVSwitch, and All That (blog.doubleword.ai)
  5. 5. The Anatomy of an Instruction Pipeline Hazard (hiraditya.github.io)
  6. 6. Width vs. Depth: Speculating on the Margin (blog.doubleword.ai)
  7. 7. Pushing memory bound CUDA kernels past the speed of light with data compression (fergusfinn.com)
  8. 8. Speculative KV coding: ~4× losslessly compressed KV cache using a small model (fergusfinn.com)
  9. 9. 70x faster cold(ish) starts for SGLang (fergusfinn.com)
  10. 10. LLM powered data structures: A lock-free binary search tree (fergusfinn.com)

Keyboard Shortcuts

j
Next item
k
Previous item
o / Enter
Open selected item
?
Show this help
Esc
Close modal / clear selection