Go & performance infrastructure
Share on
Skills
- Go
- gRPC
- WebSocket
- Docker
- PostgreSQL
- Prometheus
- Grafana
Go services that hold under load: event pipelines, Kafka and NATS migrations, concurrency correctness, and the profiling to prove it.
Throughput problems rarely announce themselves. The pipeline holds at 3,000 requests per second per core and stops scaling long before you need it to. A mutex held one line too long costs you a week of debugging. Evil Martians write Go for the layer where those numbers decide whether your product stays up.
Where our Go work sits
Cluster operations, Terraform, and CI belong to our platform engineering practice, and Rails-level tuning belongs to performance & scale. We take the Go: the services running on your cluster, the pipelines feeding them, and the binaries you ship to customers.
Teams hire Martians as Go (Golang) consultants three ways: a concurrency and performance audit on an existing service, an embedded Go engineer inside the team, or a scoped build like a broker migration or a new CLI.
- Event pipelines and brokers: Kafka and NATS topologies, partitioning, consumer groups, exactly-once semantics (EOS), and analytics ingestion into ClickHouse.
- Go services and APIs: gRPC and HTTP services, protobuf contracts, code generation, and the integration tests that keep them honest. These ship as containers, run on your cluster, and come with the Helm charts and Prometheus metrics our own Go servers already carry.
- Concurrency correctness: mutex and channel review, data races under
-racein CI, goroutine leak and deadlock detection, context cancellation that actually propagates, and concurrent tests that don’t go flaky. - Profiling and benchmarking: pprof,
runtime/trace, allocation and GC tuning, live diagnostic logging, and load tests built on xk6-cable, our Go k6 extension, wired to Prometheus and Grafana. - Go CLIs and on-premise binaries: single-binary tools your customers run inside private networks, cross-compiled for every platform they run.
- Embedded Go engineers: a senior Martian inside your team, shipping into your repo and leaving your engineers able to keep going.
Migrating an event pipeline from NATS to Kafka at 30k RPS, with zero outages
Wallarm is an API security platform protecting 20,000+ applications and APIs. Their traffic had outgrown the NATS-based pipeline carrying it. We migrated their core event pipeline from NATS to Kafka in two months with zero outages: dozens of topics, roughly 30k RPS on the busiest one, at 2,500 to 3,000 RPS per vCPU. The services on both ends of that pipeline are Go microservices, and we worked inside them for the whole migration, from the first duplicate event to the last cutover.
Duplicate events were corrupting their ClickHouse analytics. We reached exactly-once semantics using materialized views tuned to filter duplicates while preserving the throughput the system needed to keep up with incoming traffic. The pipeline is in production, handles peak loads of 30k RPS, and has headroom left.
Go infrastructure we maintain in the open
We don’t learn Go on your codebase. We maintain Go infrastructure that other companies run in production.
- imgproxy: a standalone Go image processing server, 32M Docker pulls and 10.5K GitHub stars. It’s Go over libvips through cgo, calling Quantizr, our Rust quantizer, across the C boundary, so much of the hard work is memory ownership rather than image math. Photobucket, with 35M+ monthly active users, runs it.
- AnyCable: a Go real-time server, 9M Docker pulls. One deployment serves 10,000 listeners on a single AWS c6g.2xlarge at 6% CPU.
- Lefthook: a polyglot Git hooks manager in Go, 1.2M weekly npm downloads and 7.7K stars. It’s a single Go binary cross-compiled for every platform its users run and installed through npm, Homebrew, and RubyGems.
- Overmind: a Procfile-based process manager in Go, 3.5K stars.
That work is where our Go writing comes from, including the mutex contention bug we tracked down with pprof, the error handling model we settled on after years of imgproxy, and streaming diagnostic logs out of a live Go app without restarting it.
Concurrency bugs, found before production
The hardest bugs to find are the ones a single line fixes. We encode the rule in a linter so nobody has to remember it.
At GopherCon 2026 in Seattle, Martians gave two talks on exactly this problem. Vladimir Dementyev presented Unlocking Real-World Go Mutex Usage Patterns by Writing a Mutex Linter, built after losing days to a missing Unlock(). Alexander Baygeldin presented (Sync)testing Concurrent Code with Confidence, on using Go’s testing/synctest package to make concurrent tests idiomatic, reliable, and fast instead of hacky, flaky, or slow. We write against the current release, which is why testing/synctest is in our tests and not just in a talk.
We bring the same discipline to client code: coverage from black-box tests and benchmarks rather than mock-heavy white-box runs, race detection in CI, and static analysis tuned to the mistakes your codebase actually makes.
Shipping into your open source repo
Plenty of our Go work lands in public repositories under someone else’s name.
We’ve been Teleport’s partner since 2020, through a $1.1B valuation. We built their access plugin ecosystem in Go, shipping into gravitational/teleport-plugins with tests, specs, and enterprise install guides, and gave API feedback while they migrated from REST to gRPC. Two more Go repositories we authored live in their organization: protoc-gen-terraform, which generates Terraform provider schemas from protobuf definitions, and protobuf-as, a Protobuf AssemblyScript compiler.
Protobuf is the contract we work from: schemas that stay backward compatible, generated clients in the languages your customers use, and streaming endpoints that don’t fall apart under reconnects.
These plugins are currently the #1 selling point for our Teleport product.

Sasha Klizhentas
CTO at Teleport
For NATS, the bottleneck sat on the client side, not in the Go server. We profiled the Ruby SDK, found that a thread per subscription was starving the scheduler at tens of thousands of mostly idle threads, and replaced it with a thread pool, making it more than 3x faster under high load.
This was a major enhancement to the client so I appreciate the practical solution that was implemented to boost the client’s performance and scalability.

Wally Quevedo
Synadia Communications, NATS core maintainer
Adding Go to a codebase that isn’t Go
Most products reach Go from somewhere else, and the interesting question is which parts move.
For Doximity, the digital platform used by over 80% of US physicians, we added Go-powered real-time features to their Rails monolith rather than rewriting it. Business logic stayed in Rails. The high-throughput path moved to Go. We’ve written up the general version of that call: keep the framework for auth, billing, and CRUD, and delegate the heavy work to Go, C, or Rust.
For Photobucket, we embedded a developer who shipped a new Go feature on a tight deadline and left their team able to release Go features on their own. For Keygen, we built Keygen Relay, an offline-first on-premise CLI server in Go for customers whose infrastructure never touches the cloud.
What it takes to start
A 2-week sprint is $14,000. A senior Martian engineer works alongside your team, profiles the path that’s actually hot, and ships the first fix rather than a report.
Our whole backend team writes Go, alongside Ruby and TypeScript, and reaches for it when the workload calls for it. The engineers who maintain imgproxy, AnyCable, and Lefthook are the same ones who staff client projects, so scaling an engagement up doesn’t mean hunting for a Go specialist, and a Martian rotating off doesn’t take your Go knowledge with them.
From you: access to the repository and your production metrics, plus time with the engineers who own the service.
From us: a senior Martian for two weeks at $14,000. Extendable at $7,000 per week.

Irina Nazarova CEO at Evil Martians