To get the best candidate experience, please consider applying for a maximum of 3 roles within 12 months to ensure you are not duplicating efforts.
Job Category
Software Engineering
Job Details
Platform Engineering — Cloud Infrastructure
Overview of the Role
SMTS role is part of our Platform Engineering team within the Cloud Infrastructure organization. Platform Engineering is made up of platform engineers, SREs, and DevOps specialists who design, build, and operate the internal developer platform powering hundreds of Kubernetes clusters across AWS, Azure, GCP, and OCI. Whether we are automating cluster lifecycle management, hardening our GitOps delivery pipelines, or embedding AI into our daily engineering and infrastructure operations workflows, we strive to give every product team a fast, secure, and reliable path to production.
We are looking for a Senior Member of Technical Staff to accelerate the evolution of our multi-cloud platform and operational resilience. In addition to building and operating core platform services in Go and Python, you’ll get a chance to shape our GitOps, continuous deployment, and automated operations strategy, drive automation at fleet scale, and act as an AI amplifier for the team — multiplying engineering and operational output by championing agentic, AI-assisted engineering and automated remediation practices in everything we build and operate.
What You’ll Actually Be Doing
Success will be measured by the reliability, scalability, operational resilience, and developer experience of the platform services you build and operate across our multi-cloud Kubernetes fleet.
Design, build, and operate platform services and infrastructure automation in Go and Python that manage Kubernetes clusters at scale across AWS, Azure, GCP, and OCI.
Build and improve continuous deployment pipelines using GitOps tooling (Flux, Argo CD) and infrastructure-as-code frameworks (Pulumi, Terraform) to enable safe, repeatable, fully automated infrastructure releases.
Drive fleet-wide initiatives such as cluster lifecycle automation, upgrade orchestration, policy enforcement, observability improvements, and disaster recovery readiness.
Partner with SRE, security, and product engineering teams to understand operational bottlenecks and deliver self-service platform capabilities that reduce toil and accelerate delivery.
Be an AI amplifier: embed AI tooling into everyday engineering and operations workflows — using agentic coding assistants (e.g., Claude Code) for development, code review, infrastructure automation, and intelligent operational runbooks — and help raise the team’s AI fluency and force-multiply its output.
Bring an agentic mindset to platform and operational problems: identify toil and repetitive operational work, and design agent-driven or AI-augmented automations that let the platform operate, observe, and heal itself with minimal human intervention.
Participate in design reviews, write clear technical documentation, operational playbooks, and RFCs, and mentor junior engineers on platform, operations, and cloud-native best practices.
Contribute to on-call rotations and continuously improve the operational posture, monitoring, and incident response capability of the platform through automation and post-incident learning.
You’re Our Person If…
5+ years of professional experience in cloud infrastructure engineering, operations, and continuous deployment.
Strong, hands-on Kubernetes experience — operating, troubleshooting, and automating clusters in high-scale production environments (controllers/operators, networking, scaling, upgrades).
Good programming skills in Golang and/or Python, with experience building production-grade services, CLIs, or infrastructure automation tooling.
Multi-cloud infrastructure and operations experience, primarily AWS, with working knowledge of core compute, networking, IAM, storage, and managed Kubernetes services.
Hands-on GitOps experience with Flux or Argo CD, and infrastructure-as-code experience with Pulumi (or comparable tooling such as Terraform).
Demonstrated AIOps and automation fluency — you actively leverage AI tools and agentic workflows to build closed-loop automations, accelerate root-cause analysis, and design intelligent self-healing systems.
An agentic operations mindset — you don't just use AI as a chat box; you know how to delegate complex engineering and operational tasks to AI agents, orchestrate multi-agent workflows for incident response, and use AI to drastically reduce MTTR (Mean Time to Resolution).
Strong communication and collaboration skills, with the ability to work across teams, manage incidents, and drive technical initiatives independently.
Even Better If…
Hands-on experience with Claude Code or other agentic coding and operations tools — building agentic workflows, custom commands, or MCP integrations is a strong plus; a track record of being an AI amplifier on prior infrastructure teams stands out even more.
Experience with GCP, OCI, or Azure beyond AWS, especially managing and operating Kubernetes (GKE, OKE, AKS) across providers.
Experience operating in compliance-driven environments (FedRAMP, SOC 2) with strong security, reliability engineering (SRE), and change-management practices.
Familiarity with policy-as-code (Kyverno, OPA/Gatekeeper), supply-chain security (cosign, SBOM), or infrastructure observability frameworks (Prometheus, OpenTelemetry).
Experience with internal developer platforms (IDPs), platform APIs, or developer and operator experience tooling.
Contributions to open-source cloud-native or infrastructure projects (CNCF ecosystem).