circle

Senior Site Reliability Engineer - Infra Ops

circle$152K — $205K *
US-AnywhereRemote in San Francisco, CA
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering.
  • Deep expertise in designing and operating Kubernetes clusters at scale.
  • Strong Terraform skills for creating reusable infrastructure modules and workflows.
  • Production experience in programming languages like Go, Python, or JavaScript/TypeScript.
  • Proven track record improving reliability and performance in distributed systems.
  • Strong skills in observability and troubleshooting using a variety of metrics and logs.
  • Understanding of CI/CD practices and secure rollout strategies.

Responsibilities

  • Design, build, and operate Kubernetes platforms for production services.
  • Implement infrastructure as code with reusable Terraform modules.
  • Develop backend services and tools in Go, Python, or JavaScript/TypeScript.
  • Collaborate with teams to translate workload requirements into resilient solutions.
  • Enhance production lifecycle with CI/CD and deployment automation.
  • Establish observability practices across various monitoring tools.
  • Lead incident response and perform root-cause analysis to improve reliability.

Benefits

  • Remote work flexibility with the option to work from various locations.
  • A commitment to inclusive financial practices and transparency.
  • Opportunities for mentorship and team growth.
  • Focus on innovative technology and data-driven solutions.
Full Job Description
What you'll be responsible for:

As a Senior Site Reliability Engineer on Circle's platform team, you'll design, build, and operate the secure, scalable platform infrastructure behind critical digital-assets, AI, and application workloads. You will bring an engineering mindset to production operations: writing and maintaining services and automation, developing reliable Kubernetes platforms, and using Terraform to make infrastructure repeatable, auditable, and easy to evolve.

You'll work closely with platform, product, and application engineering teams to translate workload requirements into resilient technical designs across hybrid and public-cloud environments. This role is for an experienced SRE or infrastructure engineer who enjoys solving hard distributed-systems problems, taking ownership of production outcomes, and raising the reliability, performance, security, and cost-effectiveness of the systems our customers depend on.

What you'll work on:
  • Design, build, and operate Kubernetes platforms that provide secure, highly available, and scalable foundations for critical production services across hybrid and public-cloud environments.
  • Build infrastructure as code with Terraform, creating reusable modules, safe delivery workflows, and well-governed infrastructure changes.
  • Develop backend services, internal tools, and operational automation in Go, Python, or JavaScript/TypeScript to eliminate manual work and improve the developer experience.
  • Partner with engineering and product teams to understand workload requirements and design pragmatic solutions for reliability, performance, capacity, security, and cost.
  • Improve the production lifecycle through reliable CI/CD, deployment automation, progressive delivery, and clear operational ownership.
  • Define and evolve observability practices across metrics, logs, traces, alerting, and dashboards so teams can detect issues early and troubleshoot effectively.
  • Own production reliability by participating in on-call, leading incident response, performing root-cause analysis, and driving blameless postmortems and durable corrective actions.
  • Establish and maintain reliability targets through meaningful SLIs, SLOs, error budgets, capacity planning, disaster-recovery testing, and resilience improvements.
  • Embed security and compliance into platform operations, partnering with Security to protect infrastructure, workloads, and data while meeting applicable regulatory requirements.
  • Apply AI-assisted and data-driven operational techniques to improve signal detection, reduce alert noise, accelerate root-cause analysis, and surface opportunities for automation.
  • Raise the bar for the team through thoughtful code reviews, documentation, knowledge sharing, and mentorship.
  • Mentor and support team growth, fostering collaboration and scalability.


What you'll bring to Circle (not all required):
  • 5+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or a closely related software engineering role supporting production systems.
  • Deep, hands-on Kubernetes expertise: designing, operating, securing, and troubleshooting production clusters and containerized workloads at scale.
  • Strong Terraform experience, including authoring reusable modules, managing state and environments, and delivering infrastructure changes through reviewable, automated workflows.
  • Production software-development experience in Go, Python, or JavaScript/TypeScript, with the ability to build maintainable backend services, tooling, and automation-not only scripts.
  • Demonstrated success improving the reliability, performance, scalability, or cost efficiency of distributed systems in production.
  • Experience with cloud infrastructure and core networking concepts, including IAM, DNS, load balancing, routing, service networking, and secure connectivity.
  • Strong observability and troubleshooting skills using metrics, logs, traces, alerting, and incident data to diagnose complex systems.
  • Experience defining and operating against SLIs, SLOs, error budgets, incident-management processes, postmortems, and disaster-recovery practices.
  • Familiarity with CI/CD, GitOps or deployment automation, and safe rollout strategies such as canary or blue-green deployments.
  • A security-minded approach to infrastructure and a track record of partnering effectively with Security and engineering teams in regulated or high-availability environments.
  • Clear written and verbal communication, strong ownership, and the judgment to balance speed, risk, and operational excellence.
  • Experience applying AI-assisted tooling to engineering or operations workflows is a plus.


Circle is on a mission to create an inclusive financial future, with transparency at our core. We consider a wide variety of elements when crafting our compensation ranges and total compensation packages.

Starting pay is determined by various factors, including but not limited to: relevant experience, skill set, qualifications, and other business and organizational needs. Please note that compensation ranges may differ for candidates in other locations.

Base Pay Range: $152,500 - $205,000

#LI-Remote

About circle

Circle is a global financial technology firm that enables businesses of all sizes to harness the power of digital currency and public blockchains for payments, commerce and financial applications worldwide. Circle's platform has supported over 100 million transactions worth tens of billions of dollars, with nearly 10 million retail customers, over a thousand businesses, while storing and securing more than $5 billion in digital currency assets. Circle is also a principal developer of USD Coin (USDC), which together with Coinbase and the Centre Consortium oversees the standards and protocol for what has become the fastest growing, regulated, fully-reserved stablecoin. USDC now stands at well over $10 billion market cap and is adding nearly $300 million net new digital dollars in circulation every week. Today, Circle's transactional services, business accounts, and platform APIs are giving rise to a new generation of financial services and commerce applications that hold the promise of raising global economic prosperity for all through programmable internet commerce.
Learn more about circle
Size
300 employees
Industry
Founded
2013

Similar Jobs

More Jobs at circle

More Information Technology Jobs

Find similar Senior Site Reliability Engineer - Infra Ops jobs: