Train a LLM from Scratch
github.com
1 thread
Curious — how did you handle training stability early on? Was convergence an issue without heavy tuning?
Curious — how did you handle training stability early on? Was convergence an issue without heavy tuning?