Site Reliability Engineer II

NationsBenefits, LLC

$100K — $130K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 3-5 years in SRE, DevOps, Infrastructure Engineering, or Production Support.
  • Hands-on experience with incident response and troubleshooting.
  • Familiarity with Datadog or similar monitoring tools.
  • Strong background in Kubernetes and container management.
  • Experience with SQL, MySQL, or NoSQL databases.
  • Effective in high-volume production environments.
  • Strong analytical and communication skills.

Responsibilities

  • Respond to production incidents, triaging and resolving issues.
  • Monitor alerts using Datadog and other tools to ensure system health.
  • Perform root cause analysis and escalate incidents based on SLAs.
  • Collaborate with senior engineers on complex issues resolution.
  • Continuously monitor application and infrastructure performance.
  • Automate operational tasks using Python, PowerShell, Bash, or Java.
  • Maintain documentation for incidents and ensure operational compliance.

Benefits

  • Work on technology that helps millions of healthcare members.
  • Be part of a supportive and innovative engineering culture.
  • Access to modern cloud-native technologies at enterprise scale.
  • Unlimited Paid Time Off (PTO).
  • Opportunities for career growth and development.
  • Contribute to high-impact projects with global teams.
  • Achieve a healthy work-life balance while supporting critical systems.
Full Job Description
Location: Remote (US-Based Candidates Only)

Site Reliability Engineer II (SRE)


Position Overview

We are seeking a Site Reliability Engineer II (SRE) to join our growing Site Reliability Engineering team. In this role, you will help ensure the availability, reliability, and performance of our production platforms by monitoring systems, responding to incidents, troubleshooting infrastructure issues, and driving automation initiatives.

You will collaborate closely with Development, DevSecOps, and Engineering teams to maintain highly available cloud-native applications while supporting mission-critical healthcare and fintech services.

This position is ideal for someone who enjoys solving production challenges, improving operational efficiency, and working in a fast-paced environment.

Key Responsibilities

Incident Management
  • Serve as the first responder for production incidents by identifying, triaging, and resolving issues.
  • Monitor and respond to alerts generated by Datadog and other monitoring platforms.
  • Perform initial root cause analysis and escalate incidents according to defined SLAs.
  • Communicate incident status and resolution updates to internal stakeholders.
  • Partner with senior engineers to resolve complex production issues.

Monitoring & Platform Reliability
  • Continuously monitor application health, infrastructure performance, and system availability.
  • Configure and optimize monitoring dashboards and alert thresholds.
  • Troubleshoot Kubernetes environments, including pod failures, deployment rollbacks, and log analysis.
  • Support containerized applications running in Kubernetes and Docker environments.

Production Support
  • Participate in a weekday "Follow-the-Sun" production support model with global engineering teams.
  • Participate in an on-call rotation for critical production systems as needed.
  • Help maintain high availability and system uptime.

Automation & Continuous Improvement
  • Develop automation scripts and operational tools using one or more of the following:
    • Python
    • PowerShell
    • Bash
    • C#
    • Java
  • Support CI/CD pipeline monitoring and deployment reliability.
  • Contribute to self-healing solutions and automation initiatives to reduce manual operational tasks.

Collaboration
  • Work closely with Software Engineers, DevSecOps, Infrastructure, and Platform teams.
  • Recommend improvements to monitoring, tooling, and operational processes.
  • Collaborate effectively with globally distributed engineering teams.

Documentation & Compliance
  • Maintain accurate documentation for incidents, troubleshooting procedures, and post-incident reviews.
  • Ensure operational processes align with industry security and compliance standards, including HIPAA, PCI DSS, SOC 2, ISO 27001, and HITRUST.

Required Qualifications
  • 3-5 years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or Production Support.
  • Hands-on experience with production incident response, troubleshooting, and escalation.
  • Experience with Datadog or similar monitoring and observability platforms.
  • Strong experience with Kubernetes, including monitoring, troubleshooting, and workload management.
  • Experience with Docker or other container technologies.
  • Working knowledge of SQL, MySQL, or NoSQL databases.
  • Ability to work effectively in high-volume, mission-critical production environments.
  • Strong analytical, troubleshooting, and problem-solving skills.
  • Excellent written and verbal communication skills.
  • Willingness to work weekday shifts as part of a global Follow-the-Sun support model.

Preferred Qualifications
  • Experience with cloud platforms such as Microsoft Azure, AWS, or Google Cloud Platform (GCP).
  • Familiarity with CI/CD pipelines and deployment automation.
  • Experience with Helm Charts and Kubernetes deployments.
  • Knowledge of ITIL principles and Agile methodologies.
  • Experience supporting regulated environments such as healthcare or fintech.
  • Scripting or programming experience in Python, PowerShell, Bash, Java, or C#.

Why Join NationsBenefits?
  • Work on technology that positively impacts millions of healthcare members.
  • Join a collaborative, innovative, and supportive engineering culture.
  • Exposure to modern cloud-native technologies and enterprise-scale infrastructure.
  • Competitive compensation and comprehensive benefits.
  • Unlimited Paid Time Off (PTO).
  • Opportunities for career growth and professional development.
  • Work with talented global engineering teams on challenging, high-impact projects.
  • Maintain a healthy work-life balance while contributing to mission-critical platforms.

Similar Jobs

More Jobs at NationsBenefits, LLC

  • Data Engineer II
    $90K — $120K *
    Plantation, FL 33388 (Broward County)
    Healthcare
    In-Person
  • Site Reliability Engineer II
    $100K — $130K *
    Plantation, FL 33388 (Broward County)
    Information Technology
    In-Person
  • Senior Java Cloud Engineer
    $100K — $130K *
    Plantation, FL 33388 (Broward County)
    Finance & Insurance
    In-Person
  • Lead Java Cloud Engineer
    $120K — $150K *
    Plantation, FL 33388 (Broward County)
    Finance & Insurance
    In-Person
  • Regional Manager
    $90K — $120K *
    Phoenix, AZ 85032 (Maricopa County)
    Healthcare
    In-Person

More Information Technology Jobs

Find similar Site Reliability Engineer II jobs: