You will design systems that make AI fast, safe, and scalable at global scale.
What you will work on- Large scale untrusted code execution and sandboxing
- Massively parallel AI agent runtime and scheduling systems
- Multi cloud and customer VPC deployment architecture
- Right now we are looking for two main roles:Highly reliable distributed systems with strict security and data guarantees
- Deep observability across AI workflows and infrastructure
- Enterprise grade integration and metadata platforms
What we look for- Strong background in distributed systems, scalability, multi cloud architecture, and security
- Experience with operating systems, containers, K8s, cloud networking, and event driven runtimes like Knative or KEDA
- Passion for building simple, elegant solutions to hard systems problems
- Experience designing, building, and operating large scale infrastructure
- Strong zero to one mindset
Why this role matters- You will define how AI systems run in production safely
- You will work on problems where traditional infrastructure patterns break
- You will define the operating system for future AI development