Medallia

Senior Site Reliability Engineer, GovCloud

Medallia$128K — $190K *
Technical Services
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • US residency and eligibility to work without sponsorship required.
  • Bachelor's degree in Computer Science or related field, or equivalent experience.
  • 5+ years in Site Reliability Engineering, platform engineering, DevOps, or similar roles.
  • Production experience with Kubernetes and AWS cloud services.
  • Proficiency with Terraform or similar infrastructure-as-code tools.
  • Experience with incident management and troubleshooting production systems.
  • Strong programming skills in Python and/or Go for automation.

Responsibilities

  • Build and enhance secure cloud infrastructure on AWS.
  • Design AWS cloud networking including VPCs and load balancing.
  • Manage production PostgreSQL databases and optimize performance.
  • Monitor production systems and address incidents proactively.
  • Develop Infrastructure-as-Code workflows with Terraform.
  • Enhance observability across systems and refine operational runbooks.
  • Collaborate with engineers and security teams on production changes.

Benefits

  • Comprehensive health and wellness benefits, including medical, dental, and vision coverage.
  • 401(k) retirement plan with company matching.
  • Short-term and long-term disability insurance.
  • Life and AD&D insurance.
  • Generous paid parental leave and paid holidays.
  • Opportunities for professional development and participation in diverse workplace initiatives.
Full Job Description
Overview

The Role and Team

We are growing our GovCloud team and looking for a Senior Site Reliability Engineer to help operate and improve Medallia's US public-sector cloud platform. You will support federal agencies and other regulated customers in a highly available, secure, and compliant environment built on AWS GovCloud and Kubernetes.

This is a hybrid role based near Tysons, Virginia, with regular in-office collaboration and remote flexibility. You will be a hands-on engineer who owns production systems, partners closely with engineering and security teams, and helps keep the platform reliable as we scale.

Responsibilities

  • Build, operate, and improve highly available, secure cloud infrastructure on AWS, including networking, access management, Kubernetes clusters, DNS, certificates, and shared platform services.
  • Design and operate AWS cloud networking end-to-end - VPC architecture, subnetting and routing, security groups/NACLs, VPC endpoints/PrivateLink, Transit Gateway, load balancing, and DNS - for secure, segmented, highly available connectivity.
  • Operate and tune production PostgreSQL - high availability and replication, backups and recovery, query and performance optimization, version upgrades, and capacity planning - as part of the platform's data tier.
  • Monitor production systems, respond to incidents, and drive fixes that improve reliability and reduce repeat issues.
  • Develop and maintain Infrastructure-as-Code (primarily Terraform) and Kubernetes deployment workflows using Git, CI/CD, and GitOps practices.
  • Improve observability across metrics, logs, and uptime monitoring; help tune alerts and operational runbooks.
  • Work with software engineering, security, and release management teams to deploy changes safely and resolve production issues.
  • Contribute to platform upgrades, security patching, and compliance-driven maintenance in a regulated cloud environment.
  • Participate in an on-call rotation for production support.
  • Document systems and operational procedures clearly.
  • Use AI-assisted tooling responsibly, with attention to security, privacy, and customer data boundaries.

Candidates based in the Tysons vicinity will be prioritized as this role is Hybrid, 3 days per week onsite.

Qualifications

Minimum Qualifications
  • Must reside in the United States and be legally authorized to work in the US without sponsorship.
  • Bachelor's degree or equivalent experience in Computer Science or a related field.
  • 5+ years of experience in Site Reliability Engineering, platform engineering, DevOps, or related production infrastructure roles.
  • Production experience with:
    • Kubernetes
    • Core services (IAM, compute, object storage, encryption/key management) and and cloud networking on AWS (strongly preferred), Google Cloud (GCP), Azure, or a similar public cloud platform.
    • Terraform or comparable infrastructure-as-code tools
    • Git and CI/CD pipelines
    • Linux and foundational systems concepts (networking, DNS, TLS/certificates)
    • PostgreSQL (or comparable relational databases) in production - replication, backups, and performance tuning
  • Programming and Automation: Proficiency in Python and/or Go experience to build automation scripts, operational tooling, and infrastructure services.
  • Incident & Change Management: Experience troubleshooting production incidents, conducting root-cause analysis(RCA), and following change management processes.
  • Experience participating in a production on-call rotation.
  • Experience troubleshooting complex technical issues and writing clear documentation, runbooks and incident post-mortems.

Preferred Qualifications
  • Experience operating in FedRAMP, AWS GovCloud, or other regulated or compliance-heavy cloud environments.
  • Familiarity with security and compliance practices such as FIPS and vulnerability management.
  • Experience with observability and logging platforms in enterprise production environments.
  • Deep operational expertise with PostgreSQL (HA/replication, tuning, backup and recovery); familiarity with Redis and Kafka.
  • Experience supporting federal agencies or public-sector customers.
  • Experience with tools such as Jenkins, Argo CD, and GitHub Enterprise.
  • Strong collaboration skills and willingness to learn in a compliance-driven environment.


Medallia is committed to equal pay and transparency. The annual base salary range for this position is $128,500 - $190,000. Please note that the salary range information provided is a general guideline and combines all of the distinct labor markets within the US. It is uncommon for an individual to be hired at or near the top of the range for their role and compensation decisions are dependent on a variety of factors. Medallia considers factors such as (but not limited to) scope and responsibilities of the position, candidate's work experience, candidate's work location, education/training, key skills, internal peer equity, external market data, as well as, market and business considerations when making compensation decisions.

Medallia also offers competitive health and wellness benefits, including but not limited to medical, dental, vision, 401(k), short-term and long-term disability, life and AD&D insurance, statutory leaves, paid parental leave, and paid holidays. Benefits and eligibility may vary by location and role.

At Medallia, we celebrate diversity and recognize the value it brings to our customers and employees.

About Medallia

Medallia is a software company that provides customer experience management solutions. The company was founded in 2001 by Borge Hald and Amy Pressman and is headquartered in San Francisco, California. Medallia's software allows businesses to collect and analyze customer feedback across multiple channels, including email, social media, and mobile. The company's clients include some of the world's largest brands, such as Hilton, Delta Air Lines, and Mercedes-Benz. Medallia went public in 2019 and is traded on the New York Stock Exchange under the ticker symbol MDLA.
Learn more about Medallia
Size
2,037 employees
Market Cap
$5.3 billion
Industry
Net Income
-$148.6 million
Founded
2001
Revenue
$477.2 million
NASDAQ

Similar Jobs

More Jobs at Medallia

More Technical Services Jobs

Find similar Senior Site Reliability Engineer, GovCloud jobs: