medicis123
- Karma
- 7
- Created
- 8 years ago
Recent Submissions
- 1. ▲ New Inference Server for DGX Spark: large model C4:55-90 tok/s no spec decode
- 2. ▲ Show HN: Stop GPU pods placement getting bottlenecked by reserved VRAM
- 3. ▲ A New Approach to GPU Sharing: Deterministic, SLA-Based GPU Kernel Scheduling
- 4. ▲ Show HN: Disaggregating GPU compute from CPU in ML job execution to scale GPUs (woolyai.com)
- 5. ▲ Show HN: Run PyTorch on CPU boxes, offload kernels to remote GPUs
- 6. ▲ Running Nvidia CUDA PyTorch container project/pipelines on AMD with no changes
- 7. ▲ GPU-accelerated code on CPU-only environments -Remote GPU Kernel Execution (youtube.com)
- 8. ▲ Sharing base model in GPU VRAM across multiple inference stack process [video] (youtube.com)
- 9. ▲ Sharing actual GPU core and VRAM utilization metrics for query on 10 LLM models (woolyai.com)
- 10. ▲ Show HN: WoolyAI-CUDA Abstraction Layer to Decouple Kernel Shader Exec on GPU (woolyai.com)