Last year we proposed different tests that studied single tasks. We now think that studying behavior on new tasks better captures what we want from foundation models: tools for new problems. It's what separates Newton's laws from Kepler's predictions.
