If you train a model for too long, it may overfit it's training data. Not surprising, this has been know for like forever. But did you know you can detect the signatures of overfitting in the layer weight matrices directly, without needing access to any data (train or test) ? In our recent paper (with hari kishan prakash ), ๐๐๐ญ๐-๐๐ญ๐๐ ๐ ๐๐๐ง๐๐ซ๐๐ฅ๐ข๐ณ๐๐ญ๐ข๐จ๐ง ๐๐จ๐ฅ๐ฅ๐๐ฉ๐ฌ๐ ๐ข๐ง ๐๐ซ๐จ๐ค๐ค๐ข๐ง๐ : ๐๐๐ญ๐๐๐ญ๐ข๐ง๐ ๐๐ง๐ญ๐ข-๐๐ซ๐จ๐ค๐ค๐ข๐ง๐ ๐ฐ๐ข๐ญ๐ก ๐๐๐ข๐ ๐ก๐ญ๐๐๐ญ๐๐ก๐๐ซ, we show this explicitly in 2 different classic grokking experiments. And the overfitting we see is very different from what has been seen before! ๐ paper: arxiv.org/abs/2602.02859 ๐ ๐๐ก๐๐ญ ๐ญ๐ก๐ข๐ฌ ๐ฉ๐ฅ๐จ๐ญ ๐ฌ๐ก๐จ๐ฐ๐ฌ: Grokking โ Stability โ Anti-Grokking This figure below tracks training accuracy (red), test accuracy (purple), and WeightWatcher correlation traps (blue) while training for very long times. ๐๐ก๐๐ฌ๐ ๐ โ Memorization (pre-grokking) Training accuracy rises rapidly while test accuracy remains low. The model is fitting the training data without extracting the underlying structure. Correlation traps are minimal and largely uninformative. ๐๐ก๐๐ฌ๐ ๐ โ Grokking Test accuracy suddenly jumps to match training accuracy. The model transitions from memorization to true generalization. Correlation traps remain near zero, indicating stable, well-conditioned internal representations. ๐๐ก๐๐ฌ๐ ๐ โ Late-stage instability (anti-grokking) Despite perfect training accuracy, test accuracy degrades over time. At the same time, correlation traps increase sharply and spread. Generalization collapses after it was achieved. The model is overfit ๐๐๐ฒ ๐ญ๐๐ค๐๐๐ฐ๐๐ฒ So what ? It turns out, a lot of open-source LLMs, like OpenAI's GPT OSS 20B and 120B , show the exact same signatures! And a ton of them If you are training or fine-tuning your own models, watch out! You might be overfitting your data, even if you are following current NN best practices. Want to learn more? Check out the WeightWatcher project ๐ weightwatcher.ai ๐๐๐ข๐ ๐ก๐ญ๐๐๐ญ๐๐ก๐๐ซ ๐ข๐ฌ ๐ ๐จ๐ง๐-๐จ๐-๐-๐ค๐ข๐ง๐ ๐ฆ๐ฎ๐ฌ๐ญ-๐ก๐๐ฏ๐ ๐ญ๐จ๐จ๐ฅ ๐๐จ๐ซ ๐๐ง๐ฒ๐จ๐ง๐ ๐ญ๐ซ๐๐ข๐ง๐ข๐ง๐ , ๐๐๐ฉ๐ฅ๐จ๐ฒ๐ข๐ง๐ , ๐จ๐ซ ๐ฆ๐จ๐ง๐ข๐ญ๐จ๐ซ๐ข๐ง๐ ๐๐๐๐ฉ ๐๐๐ฎ๐ซ๐๐ฅ ๐๐๐ญ๐ฐ๐จ๐ซ๐ค๐ฌ (๐๐๐๐ฌ). And you need help with AI, reach out. hashtag#TalkToChuck P.S. I'll be giving a talk at USF right here in SF (over by the GG Park) in 2 weeks on the weightwatcher project. Hope to see you there!
