About the RoleAs a Software Engineer on this team, you will build and operate the infrastructure that supports frontier research and critical research-facing systems. You will work on systems that sit close to the metal, but the role is not limited to classic operations or sysadmin work. We are looking for someone who can reason about networking, bootstrapping, Kubernetes, scalability, automation, and reliability - while also writing software to make these systems better over time.
This role is a strong fit for an independent, high-ownership engineer who enjoys reliability-heavy infrastructure work but still wants to build. You do not need to come in as a kernel expert or highly algorithmic optimization engineer, but you should be deeply curious about infrastructure, comfortable debugging complex systems, and excited to support researchers doing novel work.
We expect you to:- Build and operate reliable infrastructure for research workloads and research-facing services.
- Support and improve systems across data infrastructure, processing, crawl and ingest, caching, search, observability, and clusterwide services.
- Improve cluster bootstrapping, provisioning, automation, and deployment workflows.
- Debug issues across networking, compute, storage, orchestration, and service reliability layers.
- Build software and automation that reduce manual operational work and improve system reliability.
- Partner closely with researchers, infrastructure engineers, and service owners to understand system needs and translate them into durable solutions.
- Help evolve existing infrastructure toward more scalable, maintainable, and standard patterns.
- Take ownership of critical systems and drive work independently from problem definition through execution.
You might thrive in this role if you:- Have strong systems fundamentals and understand how infrastructure scales in practice.
- Are comfortable with Linux, networking, Kubernetes, provisioning, and distributed systems operations.
- Can write software to automate, debug, and improve infrastructure systems.
- Have a strong execution mindset and can independently drive ambiguous infrastructure work.
- Enjoy supporting a wide surface area of systems, from research tooling to platform services.
- Are pragmatic about when to build custom systems versus using existing, well-supported tools.
- Care about building reliable systems that make researchers faster and reduce operational friction.
Nice to have:- Experience with PXE boot, cluster provisioning, bare-metal infrastructure, or large-scale fleet management.
- Experience operating Kubernetes or similar orchestration systems at scale.
- Experience with infrastructure-as-code, CI/CD, observability, or deployment automation.
- Experience supporting search infrastructure, data platforms, ingest systems, or large-scale research workflows.
- Experience with Git-based workflows and internal developer tooling.
- Prior experience in environments where reliability, scale, and speed all matter.