We benchmarked coding agents on Rails for the Rails Foundation. 41 tasks on Writebook and Fizzy, 17 models, 1,770 runs, all open source.
The hard part was building a test the best models cannot max out. It worked: 35% at Stage 2. Models never reach for delegated_type or purge_later, and they are worst of all at migrations and i18n. Plenty of other surprises.
We also got to build and open source the harness, rails/lemans. I will show you how to point it at your own repo, because the best test suite is always your own codebase.




