AI research lab building RL environments to train interpretability and alignment research agents.
d_model is a fundamental AI research lab that partners with frontier AI labs to turn their models into capable interpretability and alignment researchers. The lab builds reinforcement learning environments for open-ended interpretability tasks — designed to teach agents how to do cutting-edge research rather than reward-hack. Users can contribute to training the next generation of models by participating in these environments. d_model is a product of d_model.
For people
For agents
Nothing listed yet.