mezark
- Karma
- 126
- Created
- 3 years ago
Recent Submissions
- 1. ▲ What happens when you run a CUDA kernel? (fergusfinn.com)
- 2. ▲ A running list of reasons to move to open source (whyopensource.ai)
- 3. ▲ Moe inference optimizations: 15% lower expert load by request reordering (blog.doubleword.ai)
- 4. ▲ Tensor Network Attention (mainlymatmul.com)
- 5. ▲ Redundant Information in LLM Weights (fergusfinn.com)
- 6. ▲ Tans: Precomputing RANS (fergusfinn.com)
- 7. ▲ Also-RANS: Asymmetric Numeral Systems for Entropy Coding (fergusfinn.com)
- 8. ▲ 70x faster cold(ish) starts for SGLang (fergusfinn.com)
- 9. ▲ QueueSpec – drafting speculation tokens while a request queues (blog.doubleword.ai)
- 10. ▲ ZeroDP: Just-in-Time Weight Offloading over NVLink for Data Parallelism (mainlymatmul.com)