Staff Backend Software Engineer, Developer Productivity Engineering

Tapestry

$207K — $290K *
Enterprise Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 8+ years of experience in Infrastructure, Site Reliability Engineering, DevOps, or Developer Productivity/Platform Engineering.
  • 2+ years in a Tech Lead or Staff Engineer role overseeing technical roadmaps and system architecture.
  • Expertise in Kubernetes (GKE) and multi-tenant cloud architectures, preferably on Google Cloud Platform.
  • Deep experience with Infrastructure as Code tools such as Terraform or Pulumi.
  • Proven ability to create automated build/test/release pipelines that facilitate developer self-service platforms.

Responsibilities

  • Architect and manage scalable, multi-tenant Kubernetes clusters on GCP for diverse workloads.
  • Design efficient CI/CD pipelines to enhance product development speed and reliability.
  • Implement Infrastructure as Code to ensure reproducible and secure cloud infrastructure.
  • Create isolated sandbox environments for AI tool testing and automation.
  • Mentor and guide mid- to senior-level engineers in technical practices and system design.

Benefits

  • Competitive salary and equity compensation.
  • Medical, dental, and vision coverage.
  • Generous paid time off and flexible hybrid work model.
  • 401(k) with employer contributions.
  • Opportunities for professional development.
Full Job Description
S t a f f B a c k e n d S o f t w a r e E n g i n e e r , D e v e l o p e r P r o d u c t i v i t y E n g i n e e r i n g

Software Engineering Mountain View, CA (HQ)

About the Role

We are seeking a Staff Software Engineer / Technical Lead to co-anchor the technical direction, systems architecture, and engineering standards for our Infrastructure and Developer Productivity team.

In this role, you will partner alongside another L6 Tech Lead and the Engineering Manager to set our multi-year platform roadmap. Our systems must process massive amounts of power-grid data, run heavy scientific simulations, and support emerging AI workflows. You will treat our internal infrastructure as a product-building reliable, self-service tools for our developers, automating our delivery pipelines, and creating secure environments where engineers and automated agents can build and test software quickly and safely.
How You Will Make 10X Impact
  • Core Cloud & Compute Platform: Architect and operate scalable, multi-tenant Kubernetes clusters on Google Cloud Platform (GCP). Ensure efficient compute scheduling and autoscaling across mixed hardware workloads (CPUs, GPUs, and TPUs) supporting simulation and machine learning models. Design resilient networking, IAM boundaries, and secure multi-project cloud environments.
  • Developer Velocity & CI/CD: Design fast, hermetic, and automated build, test, and release pipelines that reduce cycle times for product engineers. Provide reliable, on-demand testing environments so teams can validate changes safely before production. Build automated deployment and rollback mechanisms with clear canary verification.
  • Infrastructure as Code & Security: Drive declarative, reproducible cloud infrastructure using modern Infrastructure-as-Code (such as Terraform). Implement automated policy checks, secret management, and container vulnerability scanning into standard deployment workflows.
  • AI Tooling & Execution Sandboxes: Design secure, isolated sandbox environments that allow automated AI tools and agents to safely run tests and inspect code. Identify high-leverage opportunities to automate repetitive developer workflows using modern AI tools.
  • Technical Co-Leadership & Mentorship: Partner with your fellow L6 Tech Lead to split architectural ownership, guide system designs, and run engineering design reviews. Mentor mid-level and senior engineers (L4/L5), raising the technical bar for code reviews, testing, and system design.
  • Reliability & Observability: Define team-wide standards for metrics, logs, and distributed tracing to ensure high visibility into production health. Partner with data and ML teams to establish SLAs/SLOs, lead disaster recovery exercises, and run blameless post-mortems.
What You Should Have
  • Technical Breadth & Depth: 8+ years of production experience in Infrastructure, Site Reliability Engineering, DevOps, or Developer Productivity/Platform Engineering.
  • Technical Leadership: 2+ years serving as a formal Tech Lead or Staff Engineer directing the technical roadmap, architectural designs, and execution for a multi-pod or multi-team engineering surface.
  • Modern Cloud & Orchestration: Advanced production-grade expertise with Kubernetes (GKE), container networking, and multi-tenant cloud architectures (Google Cloud Platform preferred).
  • Infrastructure as Code (IaC): Deep architectural experience with modern declarative tools (Terraform, Pulumi, or similar) managing complex multi-environment cloud footprints.
  • CI/CD & Developer Experience Primitives: Proven track record building large-scale, automated build/test/release pipelines (e.g., Tekton, GitHub Actions, Argo Workflows, Bazel) designed around self-service internal developer platforms.
  • Platform-as-a-Product Mindset: Demonstrated ability to interview internal engineering stakeholders, quantify developer friction points, and deliver platforms that measurably increase overall deployment frequency and reduce MTTR.
Preferred Qualifications
  • AI & Agentic Workloads: Practical experience designing infrastructure, sandboxes, and execution runtimes specifically geared toward AI agents, LLM evaluations, or high-performance GPU orchestration.
  • Data Platform Adjacency: Hands-on architectural exposure to supporting large-scale data platforms (e.g., BigQuery, Spark, Kafka, Ray) or complex distributed simulation environments.
  • Alphabet Ecosystem: Working knowledge of Google-internal infrastructure primitives (Borg, Monarch, Spanner, Piper/Blaze) or experience operationalizing an X moonshot into an independent production environment.
  • Tier-1 Tech Platform Experience: Background engineering high-throughput platform tooling or developer infrastructure at companies operating at high engineering scale (e.g., Netflix, Snowflake, LinkedIn, Datadog).

What we offer

A culture that supports growth, ownership, and meaningful impact, along with:
  • Competitive salary and equity
  • Medical, dental, and vision coverage
  • Generous PTO and flexible hybrid work model
  • 401(k) with employer contribution
  • Professional development
  • The ability to work on important real-world problems within an Alphabet-backed environment

The US base salary range for this full-time position is $207,000 - $290,000 + bonuses + equity + benefits. Our salary ranges are determined by role, level, and location. Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training. Your recruiter can share more about the specific salary range for your location during the hiring process.

Please note that the compensation details listed in US role postings reflect the base salary only, and do not include bonus, or benefits.

Similar Jobs

More Jobs at Tapestry

More Enterprise Technology Jobs

Find similar Staff Backend Software Engineer, Developer Productivity Engineering jobs: