How to Post-Train

How to Post-Train

2 min read Original article ↗

A running series of tutorials, opinionated pieces, and guides around what actually goes wrong in post-training. Everything from trajectory eyeballing, rubrics and verifiers, task design, to RL environment quality. Written from years in the trenches :).

By Auriel · 5 posts

RL Fundamentals Mini-Series

Reference

Who this is for

  • Startups post-training their own models You have a training run, a harness, and a growing suspicion that something in your data is quietly wrong. These posts give some notes of what to look for.
  • Data and eval teams On writing rubrics, critics, and state optimizers a practitioner would actually use, and building verifiers that measure the thing you actually care about.
  • Anyone eyeballing trajectories How to read a trajectory, build a failure taxonomy, and tell a model failure apart from a harness failure before you burn a run on it.
  • People buying RL data What to ask for, what to reject, and the tells that a vendor’s environment was built by someone who Googled the domain for an afternoon.

New posts go out on LinkedIn and X first.