# LLM benchmarking

> Evil Martians build **LLM benchmarking** on your own repository: real tasks from your git history graded by tests, with costs.

- URL: https://evilmartians.com/services/llm-benchmarking
- Skills: LLMs, Agentic development

---

---

*Book a free 30-minute consultation to scope an eval suite for your codebase.* [Contact Evil Martians](https://evilmartians.com/contact-us)

---

## Evil Martians method

We pull tasks out of your git history and turn each one into a test that says pass or fail. Agents work in a sandbox, where they can't reach the network or edit the grader. Every model goes through the same suite, and we report on successes and costs. You run test suite on models as they come out.

The harness we built for this, [lemans](https://github.com/rails/lemans), is open source, as is the [task corpus](https://github.com/rails/ai-evals).

Want to know which agent writes good code in your repo? [Book a free 30-min call today](#book-a-call).

## Social proof

### Rails Foundation

The Rails Foundation wanted to see whether coding agents are good at Rails. In six weeks we built the atomic tasks and [lemans](https://github.com/rails/lemans), the harness behind [Agents on Rails](https://rubyonrails.org/ai), then moved on to feature tickets that imitate how an engineer actually works.

Irina Nazarova and Vladimir Dementyev presented the results at Rails at Scale 2026: [Agents on Rails](https://evilmartians.com/events/agents-on-rails-rails-at-scale).

<blockquote class="twitter-tweet" data-dnt="true"><a href="https://twitter.com/rails/status/2087951277573488825"></a></blockquote>
<script async src="https://platform.twitter.com/widgets.js" charset="utf-8"></script>

Stay up to date with the latest results on the [Rails Foundation's agents blog](https://rubyonrails.org/category/agents).

We started with atomic tasks. The best models (Claude Fable 5.1 and Opus 5) solved 92%. GPT-5.6 Luna solved 73% and cost 83 times less. Then we moved onto 20 feature tickets written "by a product manager". The best model, GPT-6 Astra, shipped 35%. Stage three is about higher-level tasks like migrations and spanning a codebase.
