gpjt
- Karma
- 1,803
- Created
- 17 years ago
About
https://www.gilesthomas.com/Recent Submissions
- 1. ▲ I use AI on this blog (gilesthomas.com)
- 2. ▲ Why do OpenAI's GPT-2 weights beat mine? Part three: testing overtraining (gilesthomas.com)
- 3. ▲ Why do OpenAI's GPT-2 weights beat mine? Part two: the bugfix (gilesthomas.com)
- 4. ▲ Why do OpenAI's GPT-2 weights beat mine? (gilesthomas.com)
- 5. ▲ Benchmarking Qwen 3.6 35B MoE (3B active) on an RTX 3090 (gilesthomas.com)
- 6. ▲ Building intuition about LLM parameter counts (gilesthomas.com)
- 7. ▲ Poppy the training box, part 1: the beginnings (gilesthomas.com)
- 8. ▲ From bigrams to GPT-2, one component at a time (in Jax) (gilesthomas.com)
- 9. ▲ Building a Jax training loop for an LLM training run (gilesthomas.com)
- 10. ▲ Thoughts on Role Confusion (gilesthomas.com)