About the role and team.We are looking for a Senior Software Engineer to build the infrastructure that powers our data and AI platforms.
You will design and operate distributed systems that support high-throughput data processing, real-time workloads, and production AI applications. This includes evolving our Kubernetes and cloud foundations, improving the reliability and scalability of platforms such as Kafka, Spark, and Flink, and building infrastructure for AI traffic management and model serving.
This is a high-impact role for an engineer who enjoys solving complex infrastructure problems, writing production software, and giving other engineering teams reliable self-service platforms. You will work across application, data, machine learning, security, and infrastructure teams to establish the technical foundations for the company's next stage of growth.
This is a hybrid role based in
Seattle, WA.
What You'll Do- Design, build, and operate highly available data and AI infrastructure on Kubernetes and public cloud platforms.
- Develop scalable platforms for streaming, batch processing, and real-time data workloads using technologies such as Kafka, Spark, and Flink.
- Build and evolve AI infrastructure, including AI gateways, model-routing layers, traffic management, rate limiting, authentication, observability, and usage controls.
- Develop self-service capabilities that enable data, AI, and application teams to deploy and operate workloads safely and independently.
- Partner with engineering teams to translate emerging data and AI requirements into durable platform capabilities.
What We're Looking For- 5+ years in DevOps, SRE, or platform engineering, owning production systems end to end
- Strong experience designing, operating, and troubleshooting production Kubernetes environments.
- Experience building or operating distributed data infrastructure with technologies such as Kafka, Spark, or Flink OR experience developing AI infrastructure such as AI gateways or model-routing platforms.
- Hands-on experience with at least one major public cloud platform, such as AWS, Google Cloud, or Microsoft Azure.
- Strong knowledge of cloud and container networking, including DNS, load balancing, ingress, service discovery, TLS, routing, and network security.
- Proficiency in one or more of Go, Python, or Java, with experience writing maintainable production software.
- A solid understanding of distributed-systems concepts, including availability, consistency, fault tolerance, backpressure, and horizontal scalability.
- Experience operating critical infrastructure using infrastructure-as-code, automated delivery, and modern observability practices.
- Strong debugging skills and the ability to work methodically across multiple layers of a complex system.
- Clear communication skills and a track record of collaborating effectively across engineering disciplines.
- An ownership mindset: you identify important problems, drive them to resolution, and improve the underlying system rather than treating symptoms.
Bonus Points- Platform engineering experience, particularly building internal developer platforms or paved-road workflows used by multiple engineering teams.
- Hands-on experience with serving technologies such as SGLang, vLLM, or NVIDIA Triton Inference Server.
- Knowledge of GPU scheduling, batching, model parallelism, memory management, autoscaling, and inference-performance optimization.
- Experience improving the cost efficiency of large-scale data processing or AI inference workloads.
- Contributions to infrastructure, data-platform, Kubernetes, or AI-serving open-source projects.
What Success Looks Like- Delivered meaningful improvements to the scalability, reliability, or efficiency of our data and AI infrastructure.
- Reduced the operational effort required to deploy and manage data or AI workloads.
- Improved visibility into system performance, reliability, capacity, and cost.
- Established reusable platform capabilities adopted by engineering teams.
- Helped define the technical direction for our next generation of data and AI infrastructure.