5-7 years of experience in building and scaling production cloud infrastructure.
Proficient in writing and maintaining Go services in a production environment.
Strong understanding of Kubernetes architecture and debugging techniques.
Ability to investigate failures across application code, networking, and operating systems.
Experience shipping production software using AI-assisted development.
Responsibilities
Build and launch new tools, features, and cloud services, including Go services and APIs.
Optimize startup times and improve resource scheduling as demand increases.
Implement tenant isolation, access controls, and networking for customer data protection.
Automate host provisioning, upgrades, and failure recovery processes.
Collaborate with customers to enhance workflows and develop better APIs and tooling.
Benefits
Flexible working hours in an in-person role in San Francisco.
Relocation support for moving to the San Francisco office.
Regular team offsites to foster collaboration.
Company-covered work travel opportunities across Europe and the US.
Full Job Description
Help us expand what AI agents can build. Limrun runs development tools and applications as cloud services that agents can use from wherever they run. We're hiring an engineer to build and scale that platform. The role
You'll help decide what to build and own projects from design through ongoing maintenance. You'll also join the on-call rotation, investigate production incidents, and fix their causes. What you'll do
Build and launch new tools, features, and cloud services, including the Go services, APIs, and Kubernetes controllers behind them.
Reduce startup times, improve scheduling, and make better use of compute as demand grows.
Build and maintain the tenant isolation, access controls, and networking that protect customer workloads and data.
Automate host provisioning, upgrades, and recovery from failures.
Work with customers to understand their workflows and turn problems they encounter into better APIs, tooling, and product features.
What we're looking for
You've built, operated, and scaled production cloud infrastructure, with responsibility for deployment, observability, and reliability.
You've written and maintained Go services in production and can reason about concurrent code.
You understand Kubernetes architecture and controller patterns well enough to build on them and debug their behavior.
You can investigate failures across application code, networking, and operating systems.
You've shipped production software or infrastructure using AI-assisted development.
We care about your work and the responsibility you've taken for it more than titles or years of experience. Nice to have, not required
Experience with sandboxing, virtual machines, or managing Linux and macOS hosts.
Hands-on work with DNS, routing, proxies, tunnels, or private networking.
Experience building APIs, SDKs, CLIs, CI systems, or remote build tools.
Experience using or building AI agents for platform maintenance and production operations.
Familiarity with mobile development workflows using simulators and emulators.
What we offer
An in-person role in San Francisco with flexible working hours.
A San Francisco office in South Park, with relocation support.
Regular team offsites.
Company-covered work travel across Europe and the US.
Apply
Send us your resume or profile and a short note about an engineering project you're proud of. This could be:
A system or service you took from design to production.
A scaling, performance, or reliability problem you solved.
An operational improvement that reduced failures, manual work, or recovery time.
Tell us what you personally owned, the decisions you made, and what improved. If AI contributed, tell us how you used it and checked the results. Links to code, projects, or technical writing are welcome where you can share them.