kumama
- Karma
- 5
- Created
- 9 years ago
Recent Submissions
- 1. ▲ how we monitor our rl training runs (castform.com)
- 2. ▲ Designing dev onboarding for an agent-first world (castform.com)
- 3. ▲ I post-trained a model to reliably roll a die (castform.com)
- 4. ▲ Open-Weight Models Don't Need to Win (twitter.com)
- 5. ▲ Prompt caching but for RL – 7.5x speedup on long-prompt/short-response workloads (castform.com)
- 6. ▲ Pokegents: Making multi-agent coding feel like a team (castform.com)
- 7. ▲ Grpo explained: group relative policy optimization for LLM finetuning (cgft.io)
- 8. ▲ Do RL on a model with your vector db (cgft.io)
- 9. ▲ What is reinforcement learning finetuning (youtube.com)
- 10. ▲ RAG to riches: synthetic data for training RAG agents (cgft.io)