We work on measuring and optimizing model performance in fuzzy domains.
Current post-training pushes models into a narrow set of model personas. It's hard to get out of that basin, because it's hard to measure what these models lack (things like communication, truthfulness, etc.).
Understanding how to do this new kind of behavioral optimization, through post-training, elicitation, and metric design, will be critical to safely navigating the intelligence explosion. We aim to solve this problem.