OsamaJaber
- Karma
- 264
- Created
- 11 months ago
Recent Submissions
- 1. ▲ GLM 5.3 flash on AMD GPUs 670 tok/s (twitter.com)
- 2. ▲ Agent loops stop re-reading your codebase on every turn (runinfra.ai)
- 3. ▲ GLM 5.3 Flash faster and cheaper (runinfra.ai)
- 4. ▲ The fastest and cheapest GLM 5.3 Flash endpoint (runinfra.ai)
- 5. ▲ DeepSeek V4 Pro at 207 tok/s with the full 1M context, no quantization (runinfra.ai)
- 6. ▲ Fastest Inference in MENA (runinfra.ai)
- 7. ▲ DeepSeek V4 Flash at 278 tok/s, full precision, no quantization (runinfra.ai)
- 8. ▲ The fastest full-precision Nemotron 3.5 Lightning endpoint (540 tok/s) (runinfra.ai)
- 9. ▲ With software alone, one B200 beats the LPU and gets close to Cerebras (runinfra.ai)
- 10. ▲ Kimi K3 2.78T on One CPU with 8GB RAM (github.com)