Overflow
RLmulti-agentautonomous-vehiclespythonopenenv
An autonomous vehicle fleet oversight environment for OpenEnv — Meta's open framework for RL environments.
What it is
A 2D road grid with N cars. Car 0 is controlled by an LLM agent; the rest follow simple scripted driving rules. An observer detects crashes and near-misses on every step and computes rewards based on safety. The agent chooses among accelerate, brake, lane_change_left, lane_change_right, and maintain, and has to justify its decision in natural language. A hybrid reward combines rule-based safety scoring with LLM reasoning, and Python evaluation tooling benchmarks agent safety across simulated scenarios — reducing the crash/incident rate by 80%.
Why it matters
Most RL agent benchmarks reward task completion. Fleet oversight is the inverse: the agent's job is to not cause incidents while staying useful. Overflow makes that the reward function — what an LLM would actually need to do if it were the supervisor of a Waymo-style fleet rather than the driver.
Stack
PythonOpenEnvLLM Agents