PROWL: RL agents that find failures in world models
World models can improve by having reinforcement-learning agents explore simulators and games to discover adversarial trajectories and failure cases automatically
What if world models could learn by discovery? Today we’re sharing PROWL: RL agents that explore game environments, simulators, and eventually robots to discover failures in a world model. This loop of learning is fully automated!