Member of Technical Staff - Infrastructure

Favorited

$150K — $200K *
Technical Services
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 6+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering.
  • Experience managing infrastructure for large-scale systems supporting millions of users.
  • Strong expertise with cloud infrastructure, ideally Google Cloud Platform (GCP).
  • Hands-on experience with Kubernetes and container orchestration systems.
  • Experience with monitoring and observability tools (e.g., Prometheus, Grafana).
  • Strong programming experience in languages like Python, Go, or TypeScript.
  • Deep understanding of reliability engineering practices and incident management.

Responsibilities

  • Design and maintain reliable, scalable infrastructure for real-time applications.
  • Automate tools for system reliability and efficiency.
  • Develop monitoring and alerting systems for high availability.
  • Collaborate with engineering teams to enhance service performance and observability.
  • Manage incident response and incorporate learnings into system improvements.
  • Optimize infrastructure for performance and cost efficiency.
  • Scale containerized environments using technologies like Docker and Kubernetes.

Benefits

  • Unlimited PTO to promote work-life balance.
  • 401(k) plan for future investment.
  • Comprehensive health insurance for well-being support.
  • Paid company holidays for rest and rejuvenation.
Full Job Description
About the Role

We are looking for a Member of Technical Staff, Infrastructure to help ensure the reliability, scalability, and performance of the infrastructure that powers favorited's real-time platform. You will play a key role in building and maintaining systems that support high-traffic applications used by a rapidly growing global audience.

This role is ideal for someone who enjoys solving complex infrastructure challenges, improving system reliability, and building automation that allows engineering teams to move quickly and confidently.
Responsibilities
  • Design, implement, and maintain highly reliable and scalable infrastructure supporting real-time applications.
  • Build automation and tooling to improve system reliability, deployment processes, and operational efficiency.
  • Develop and maintain monitoring, logging, and alerting systems to ensure high availability and rapid incident response.
  • Partner closely with engineering teams to improve service reliability, performance, and observability.
  • Support incident response, root cause analysis, and postmortems, ensuring learnings are incorporated into system improvements.
  • Optimize infrastructure for performance, cost efficiency, and scalability.
  • Manage and scale containerized environments using Docker, Kubernetes, and related orchestration technologies.
  • Help define and enforce reliability standards, SLOs, and operational best practices across engineering teams.
  • Continuously evaluate new infrastructure tools and practices to improve system resilience and developer productivity.


What We're Looking For
  • 6+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering roles.
  • Experience managing infrastructure for large-scale systems supporting millions of users.
  • Strong expertise with cloud infrastructure, ideally Google Cloud Platform (GCP).
  • Hands-on experience with Kubernetes, container orchestration, and distributed systems.
  • Experience implementing monitoring and observability systems (Dash0, Prometheus, Grafana, Datadog, or similar).
  • Strong scripting or programming experience in languages such as Python, Go, or TypeScript.
  • Deep understanding of reliability engineering practices including SLOs, SLIs, and incident management.
  • Strong collaboration skills and ability to work cross-functionally with engineering teams.
  • Solid understanding of TCP/IP networking and emerging protocols like HTTP/3.


Nice to Have:
  • Experience supporting real-time streaming, gaming, or large-scale consumer applications.
  • Familiarity with event-driven architectures and large-scale data processing systems.
  • Experience optimizing infrastructure costs in high-growth environments.
  • Practical experience integrating AI agents into DevOps workflows.


Salary & Benefits

Compensation: $150k - $200k base salary + options.

Benefits Include:
  • Unlimited PTO to prioritize work-life balance.
  • 401(k) plan to invest in your future.
  • Comprehensive health insurance to support your well-being.
  • Paid company holidays for time to recharge.
  • Competitive salary that values your expertise and contributions.

Where You'll Work: This is a full-time, on-site position in Santa Monica.

Similar Jobs

More Jobs at Favorited

More Technical Services Jobs

Find similar Member of Technical Staff - Infrastructure jobs: