Site Reliability / Infrastructure Engineer

General Intuition

$130K — $160K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Strong fluency in Terraform for infrastructure-as-code at scale
  • Hands-on experience with Elasticsearch for user-facing features
  • Proficient with GCP services including Kubernetes, VPC, and IAM
  • Deep experience in scaling and sharding relational databases in production
  • Calm under pressure with strong incident response skills
  • Familiarity with CI/CD processes using GitHub Actions
  • Comfortable in fast-paced startup environments

Responsibilities

  • Own on-call rotation and respond to incidents
  • Drive postmortems to prevent incident recurrence
  • Work closely with engineering teams to meet infrastructure needs
  • Ensure system reliability and scalability under growth pressure
  • Communicate clearly during high-pressure situations

Benefits

  • Access to cutting-edge technology and infrastructure
  • Opportunities for professional development in a scaling environment
  • Collaborative team culture focused on innovation
  • Chance to influence major infrastructural decisions
  • Work on high-impact projects that handle billions of clips
Full Job Description
Medal's infrastructure handles billions of clips, video ingestion pipelines, and social features at a massive scale most engineers never get to touch. The work centers on reliability, incident response, scaling, and making sure our infrastructure keeps up with our growth. You'll own the on-call rotation, drive postmortems, and work directly with engineering teams to meet their infra needs. The right person probably came through startups and scale-ups, has been in the room when things broke at 2am, has scaled databases under pressure, and knows the difference between a durable fix and a patch that buys you a week. What We're Looking For • Infrastructure-as-code: Strong fluency in Terraform, with real experience owning infrastructure-as-code at scale • Elasticsearch depth: Hands-on experience running ES for user-facing features, not just as a log sink • GCP depth: You know it maybe a little too well: Kubernetes, VPC, IAM, Cloud Logging, and the managed services ecosystem • Database scaling: Deep, hands-on experience scaling and sharding relational databases (MySQL, Postgres) in production • Incident response instincts: You can work a P0 calmly, communicate clearly under pressure, and run a postmortem that prevents recurrence • CI/CD: You've worked with GitHub Actions in a production environment • Communication (crucial!): You flag issues clearly and rapidly during incidents and lead/write actionable postmortems • Experience at startups: You are comfortable in an environment of rapid growth where scaling up is a priority • Great judgment: You know the difference between a durable, sustainable fix and a patch that buys you a week Our Stack Electron, React, Redux, Styled Components & other modern web-based technologies C# and C++ for native Windows recording & more Swift for iOS, Kotlin for Android Java, Redis, RabbitMQ, Kubernetes for backend Terraform, Salt, GitHub Actions, CircleCI for IaC and CI/CD

Similar Jobs

More Information Technology Jobs

Find similar Site Reliability / Infrastructure Engineer jobs: