Evil Martians build LLM benchmarking on your own repository: real tasks from your git history graded by tests, with costs.
Evil Martians method
We pull tasks out of your git history and turn each one into a test that says pass or fail. Agents work in a sandbox, where they can’t reach the network or edit the grader. Every model goes through the same suite, and we report on successes and costs. You run test suite on models as they come out.
Want to know which agent writes good code in your repo?
Social proof
Rails Foundation
The Rails Foundation wanted to see whether coding agents are good at Rails. In six weeks we built the atomic tasks and lemans, the harness behind Agents on Rails, then moved on to feature tickets that imitate how an engineer actually works.
We started with atomic tasks. The best models (Claude Fable 5.1 and Opus 5) solved 92%. GPT-5.6 Luna solved 73% and cost 83 times less. Then we moved onto 20 feature tickets written “by a product manager”. The best model, GPT-6 Astra, shipped 35%. Stage three is about higher-level tasks like migrations and spanning a codebase.