Tatsunori Hashimoto (@tatsu_hashimoto) on X

X (formerly Twitter) ·

2 min read Original article ↗

Post

Post

Tatsunori Hashimoto on X: "We are releasing AlpacaFarm, a simulator enabling everyone to run and study the full RLHF pipeline at a fraction of the time (<24h) and cost (<$200) w/ LLM-simulated annotators. Starting w/ Alpaca, we show RLHF gives big 10+% winrate gains vs davinci003 (https://t.co/TfCWhCE55m)"

  • user avatar

    We are releasing AlpacaFarm, a simulator enabling everyone to run and study the full RLHF pipeline at a fraction of the time (<24h) and cost (<$200) w/ LLM-simulated annotators. Starting w/ Alpaca, we show RLHF gives big 10+% winrate gains vs davinci003 (crfm.stanford.edu/2023/05/22/alp…)

  • user avatar

    We find the RLHF simulator to be very accurate. The simulated annotators are close to humans in agreement rate (65 vs 66%) at 1/45th the cost, and rankings of methods trained in simulation agree with rankings of methods trained on real human feedback.

    user avatar

    Our reference RLHF implementations give substantial improvements. The best method (PPO) provides major gains over Alpaca on win rate vs davinci003 (41->55%), and we find it gives much more detailed explanations for answers.

    user avatar

    Additionally, to enable fast and reproducible evals, we define a new automatic eval for instruction following and combine many existing eval datasets. This aggregated eval set agrees very well with the simple but real human instructions from the Alpaca live demo.

    user avatar

    user avatar

    Please see our paper for other details, like how it’s necessary to emulate inter- and intra-annotator variability to build a simulator that captures important phenomena like overoptimization (tatsu-lab.github.io/alpaca_farm_pa…)

    user avatar

  • user avatar

    You may find this interesting:

    user avatar