T+00Choose
Act under partial information.
The model sees an incomplete state and commits before the full consequences are observable.
Reinforcement learning environments
THESIS
Static worlds produce static intelligence
We programmatically generate quant research tasks inside environments built from real market data. Agents use professional tools—and build their own in Bash—to make trading decisions and develop profitable strategies.
Markets do not saturate: successful trading makes them more efficient, while edges decay and regimes shift. That makes our environments a continuously harder benchmark for improving models.
HORIZON
A decision is not a moment
Trading successfully means planning ahead multiple steps and assess trade-offs between short and longterm gains
T+00Choose
The model sees an incomplete state and commits before the full consequences are observable.
T+18HCompound
Exposure, opportunity cost and every action not taken reshape the path that follows.
T+53HRevalue
A decision can remain locally correct while becoming globally expensive as conditions drift.
T+96HAdapt
Success belongs to the model that recognizes the new regime before yesterday’s behavior becomes consensus.