Site Reliability Engineer

Ad Astra Info Systems LLC

$90K — $120K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or related field preferred; equivalent experience acceptable.
  • 2+ years of Site Reliability Engineering, Systems Engineering, or DevOps experience.
  • Strong understanding of networking concepts like load balancing, DNS, IPSec, and VPNs.
  • Experience with CI/CD, Infrastructure as Code tools (e.g., GitHub, Jenkins, Terraform).
  • Proficiency with relational and NoSQL database technologies.
  • Skills in at least one scripting/programming language (Node.js, Python, etc.).
  • Experience with Docker and Kubernetes for containerization and orchestration.

Responsibilities

  • Write automation and production code to enhance system reliability.
  • Design and maintain scalable systems across cloud environments (AWS, Azure, GCP).
  • Own reliability considerations of a multi-tenant SaaS platform.
  • Enhance observability through logging, monitoring, and alerting systems.
  • Automate workflows and infrastructure provisioning bridging development and operations.
  • Monitor alerts and incidents ensuring uptime and system performance.
  • Collaborate with teams to plan capacity and drive cost reductions.

Benefits

  • 401(k) with Profit Sharing
  • Flexible Time Off
  • Office Dog
Full Job Description
Competitive Compensation & Benefits Package * 401(k) with Profit Sharing * Flexible Time Off * Office Dog!!

POSITION SUMMARY

The Site Reliability Engineer (SRE) will ensure the performance, reliability, and scalability of our systems as we continue to grow. This role bridges the gap between software development and operations, applying software engineering principles to automate, optimize, and enhance the reliability of our infrastructure and production systems. Your role includes identifying recurring failure patterns, implementing automated solutions, and continuously improving platform performance. Leveraging your intellectual curiosity and expertise in operations and development, you will also play a pivotal role in monitoring security and reliability threats, while actively advocating effective solutions.

This role spans a genuinely wide range of work, from deep automation and greenfield infrastructure projects to legacy system support and cross-team collaboration. You'll thrive here if you enjoy variety and can move between priorities without missing a beat, and if you're motivated by helping shape reliability practices as we grow toward higher availability targets.

ESSENTIAL FUNCTIONS/CORE RESPONSIBILITIES
  • Write automation and production code to improve system reliability and performance.
  • Design, build, and maintain highly available, scalable systems across cloud environments (e.g., AWS, Azure, or GCP).
  • Own reliability and scalability considerations unique to our multi-tenant SaaS platform, including tenant isolation and blast radius containment.
  • Maintain and extend logging, monitoring, and alerting systems to enhance observability and proactive incident response.
  • Bridge development and operations by automating workflows, deployments, and infrastructure provisioning.
  • Proactively monitor and respond to alerts and incidents, ensuring system uptime and performance.
  • Support security and compliance initiatives.
  • Eliminate manual toil across legacy systems (currently being migrated away from), client onboarding, and data integration, with the autonomy to build lasting automation.
  • Collaborate with engineering, product, and operations teams to capacity plan, drive cloud cost reduction, and enhance the overall reliability and efficiency of our products
  • Support production systems, including participation in on-call rotations and performing limited after-hours maintenance.
  • Lead and contribute to post-incident reviews, driving root cause analysis and long-term solutions.
  • Document reliability patterns, runbooks, and learnings to build operational maturity
  • Other duties as assigned


POSITION REQUIREMENTS
  • Bachelor's degree in Computer Science, Engineering, or related field preferred; equivalent experience in supporting distributed software systems accepted.
  • 2+ years of experience in Site Reliability Engineering, Systems Engineering, or DevOps, with strong systems/infrastructure fundamentals and comfort using modern tooling, including AI-assisted development, backed by the judgment to design sound automation, not just generate it
  • Strong understanding of networking concepts including load balancing, DNS, IPSec, and VPNs.
  • Experience with source version control, CI/CD, and Infrastructure as Code tools (e.g., GitHub, Jenkins, Terraform, CloudFormation).
  • Working knowledge of Linux operating systems.
  • Proficiency with relational or NoSQL database technologies (both preferred)
  • Proficiency in at least one scripting or programming language (Node.js, Python, Go, Bash, PowerShell, etc.).
  • Experience with containerization and orchestration (Docker, ECS, Kubernetes).
  • Familiarity with observability tools (Graylog, New Relic, Prometheus, Grafana, ELK Stack, etc.).
  • Strong collaboration, problem-solving, and communication skills.


ESSENTIAL COMPETENCIES
  • Creative Problem Solving
  • Collaborative Communication
  • Adaptability & Flexibility
  • Sense of Urgency with Quality
  • Attention to Detail
  • Technical Aptitude


ADDITIONAL PREFERRED QUALIFICATIONS
  • Expertise in git, docker, terraform, ansible and AWS
  • Experience with blue/green or canary deployment strategies and zero-downtime releases.
  • Understanding of security best practices in cloud-native environments.
  • Background in automating large-scale infrastructure management.
  • Experience working in an agile or SaaS-based environment.


HOW PERFORMANCE IS MEASURED FOR THIS ROLE
  • Meaningful contributions to the SRE high value/team stories.
  • Timely response to infrastructure alerts and ensuring system reliability.
  • Regular preventative maintenance.
  • Contribution to the overall success of the Cloud Ops team.
  • Incident Mean time to Acknowledge (MTTA) < 15 minutes.
  • Drive availability improvements to exceed 99.95% uptime.


This full time, in-office position is located in Overland Park, Kansas. Ad Astra does not pay for relocation expenses.

Similar Jobs

More Jobs at Ad Astra Info Systems LLC

More Information Technology Jobs

Find similar Site Reliability Engineer jobs: