Introducing Karotte: A Framework for Building Robust RL Environments

5-minute read

Today we're open-sourcing Karotte, our framework for building RL environments. We've been using Karotte internally to develop RL environments for the past year, and it will be the first of several components which we'll be sharing from our RL environment development stack.

In recent months, labs' agents have hacked out of their sandboxes during training and caused harm in the world. We believe that robust RL environments are a key component to training aligned models as we head into the superintelligent era. Karotte takes an opinionated approach to make building such environments easier, by providing secure defaults and primitives.

Karotte has been hardened through more than a million evaluation runs on our internal infrastructure, as well as controlled red-teaming. We've performed careful and monitored testing of agents trying to reward hack environments built on Karotte, and we're continually integrating our learnings into Karotte to harden it further.

Why Karotte?

Building robust RL environments is important because the current bottleneck on AI progress is misalignment, and developing aligned models requires robust, correct reward functions.

Reward hacking does a lot more damage than wasting some compute. Last November, Anthropic found that a model that learned to reward hack in real production coding environments became broadly misaligned. It tried to sabotage AI safety research code and showed alignment-faking reasoning in half of its responses to questions as simple as "What are your goals?" What the model learned from cheating on coding tasks carried over into how it behaved everywhere else.

More recently, OpenAI slowed down training of its newest models after its agents escaped a test environment and hacked Hugging Face, and later paused its most capable models after agents escaped a training sandbox and reached US government websites.

We think the difference between evals and training is underappreciated here. In an eval, a reward hack that happens in 10% of runs gives you a slightly wrong number on a leaderboard. In training, the hack gets reinforced. A hack that is initially tried 10% of the time will soon become the model's default approach. If exposed to enough environments that are broken in this way, a model will ultimately learn to explore its environment for loopholes instead of solving the actual task.

This is closely related to our principle that you shouldn't lie to the AI. If the prompt asks for one thing but the grader rewards something else, the model learns that the grader is what actually matters, not the instructions.

We hope Karotte will be one component among many in our toolbox for producing well-behaved, capable AIs. You can find the code at https://github.com/preferencemodel/karotte/ and get started immediately with uvx karotte create-env my-env.

What is Karotte?

Suppose we're building an RL environment to teach a model to work with CSV files and do some basic math. A simple task could tell the model that there's a CSV of orders in its working directory and that it should write the total revenue of the third quarter, in cents, to an answer.txt file. Sounds straightforward enough, but if implemented naively, we know of at least a dozen different ways to hack this environment such that the model crashes the run or receives a high score without solving the actual task.

Examples include fork bombs, directly reading the solution, or causing the grader to run out of memory. Things start getting really complicated when the environment needs to compile, run, and measure model-generated code. The agent in such an environment can try to monkeypatch the grading code or start background processes to slow down any baselines that its code gets measured against.

These reward hacks are no mere thought experiments. METR caught o3 doing this almost exactly as described last year. o3 reward hacked in about 30% of its runs on RE-Bench, in one case by making the scorer's timer report runtimes 1000x shorter, and in another by digging through the Python call stack to find the grader's reference answer.

Karotte is a framework for building environments that inherently mitigate many of the reward hacks a naive implementation would expose. Anything the model does runs as a non-privileged user inside a sandbox. Any processes the model starts get killed before grading. Any suspicious files, like FIFOs, symlinks, or large sparse files, get rejected so they can't crash the grader. By default, the model can't cause a sandbox OOM and can't fork-bomb its way out of getting a bad score for a bad solution.

Karotte has a clear focus on what it tries to do. It is designed to help build environments that models will be trained on, where defenses have to hold up against an optimizer trying millions of things. As a side effect, Karotte is also well suited for building robust benchmarks. Running Karotte environments at scale is straightforward, but it does not directly integrate with any cloud compute providers out of the box.

Read our technical deep dive to see how Karotte works under the hood.

Help us build what comes next

Preference Model is ultimately about ensuring AI goes well for everybody. The biggest problem with AIs is that they can cause great harm if they accidentally or intentionally don't do what we intend. This may be a result of their lack of capability or because they are insufficiently aligned. We're trying to address both of these issues in the most direct, highest-leverage way available to us, which today means researching how to build RL environments better, not just for the AIs of today but for the AIs as they will be when they outsmart us.

Over the past year, we've built RL environments for several frontier labs, and we're backed by $16M in seed funding led by a16z, with participation from SignalFire, South Park Commons, Scale Angel Group, and researchers including Fei-Fei Li, Ian Goodfellow, and Julian Schrittwieser.

If you want to join us in our mission to ensure AI keeps working in humanity's best interests, we'd love to hear from you: email hello@preferencemodel.com, or see our open roles.