Senior Site Reliability Engineer

Nectar Social

$130K — $180K *
Technical Services
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years in SRE, infrastructure, or platform engineering
  • Experience scaling complex production systems
  • Hands-on expertise with AWS or similar cloud infrastructure
  • Strong programming skills for automation and tools
  • Ability to thrive in fast-paced startup environments
  • Reliability-first mindset with pragmatic approach

Responsibilities

  • Oversee the reliability and scalability of production systems handling large data volumes
  • Define SLOs, SLIs, and error budgets while building observability and alerting practices
  • Lead incident response and conduct blameless postmortems for continuous improvement
  • Enhance performance, cost efficiency, and capacity planning of cloud infrastructure
  • Strengthen infrastructure-as-code and CI/CD pipelines for resilience
  • Collaborate with engineering teams to integrate reliability into system designs

Benefits

  • Competitive compensation and early equity
  • Health, vision, and dental benefits with 401(k) match
  • Clear career growth opportunities as the company expands
  • Complimentary lunch in downtown Palo Alto
  • Exposure to advanced AI tools and influence in their application
  • Work within a collaborative team focused on innovative AI marketing solutions
Full Job Description
The Role

We're looking for a Senior Site Reliability Engineer to own the reliability, scalability, and operational excellence of the production systems that power Nectar's platform. We run high-volume data ingestion pipelines and real-time AI agents on top of a fast-growing customer base, and we need a seasoned SRE to help us scale these systems safely and keep them running flawlessly.

As one of our first dedicated SREs, you'll have outsized impact and ownership. You'll define how we measure, operate, and harden our infrastructure -- establishing the reliability foundations that the rest of the engineering team builds on as we scale.
What You'll Be Doing
  • Own the reliability and scalability of our production systems as they handle rapidly growing volumes of social data and real-time AI workloads
  • Define and drive SLOs, SLIs, and error budgets, and build the observability, alerting, and on-call practices to support them
  • Lead incident response and blameless postmortems, then turn what we learn into systemic improvements that prevent recurrence
  • Improve performance, cost efficiency, and capacity planning across our cloud infrastructure as the platform scales
  • Harden our infrastructure-as-code, deployment, and CI/CD pipelines for resilience and repeatability
  • Partner with engineering teams to embed reliability into system design and raise the operational bar across the org
What We're Looking For
  • 5+ years of experience operating production systems as an SRE, infrastructure, or platform engineer
  • Experience scaling databases, data infrastructure, or complex production platforms under significant load
  • Hands-on expertise with cloud infrastructure (AWS or similar) and infrastructure-as-code tooling
  • Solid programming skills for building automation, tooling, and operational services
  • Comfortable operating in fast-moving startup environments with high ownership and autonomy
  • A reliability-first mindset balanced with pragmatism about velocity and cost
Bonus Points
  • Experience standing up or maturing an SRE practice at an early-stage or rapidly scaling company
  • Familiarity with our tech stack: AWS, Pulumi, Postgres, ClickHouse, Turbopuffer, or Temporal
  • Background in capacity planning, performance engineering, or cost optimization at scale
What We Offer
  • Competitive compensation and early equity
  • Health, vision, and dental benefits + 401(k) match
  • Clear career growth opportunities as the company scales
  • Free lunch in the heart of University Ave. in Palo Alto
  • Deep exposure to cutting-edge AI tooling and the opportunity to shape how brands use it
  • A collaborative, ambitious team defining a new category of AI-native marketing infrastructure

Similar Jobs

More Jobs at Nectar Social

More Technical Services Jobs

Find similar Senior Site Reliability Engineer jobs: