Settings

Theme

What Is RLCD? The Secret Behind Jev

di-zhang-llm.github.io

68 points by tnspacetime · 9 comments

Reader

3 threads
firejake308

> The operational signal was always relative preference. The scalar merely hid it.

Is this another Claude-ism? "X was always Y. The Z merely hid it." Or am I overcalling it?

WalterGR

RLCD, not defined in the article, is Reinforcement Learning for Calibrated Decisions.

daemonk

Yeah the calibration is really what makes it useful in practice for quick, small decisions. Asking a LLM to give scores to a problem will yield inconsistently scaled/anchored results that changes at a whim.

The blog is pretty heavy on statistics. I'll have to study it more when I have time. Is it essentially bootstrapping results to statistically normalize the answers?

  • tnspacetimeOP

    I have not studied it properly too. Good that it has both code and note though.

Keyboard Shortcuts

j
Next item
k
Previous item
o / Enter
Open selected item
?
Show this help
Esc
Close modal / clear selection