Site Reliability Engineer (SRE)

Ova Technologies

$120K — $150K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science or a related field.
  • 3-6 years of experience in SRE, DevOps, or similar roles.
  • Strong proficiency in Linux system administration.
  • Experience with Docker and Kubernetes.
  • Scripting skills in Python, Go, Bash, or Shell.
  • Familiarity with CI/CD tools like Jenkins or GitHub Actions.
  • Understanding of networking concepts (TCP/IP, DNS, HTTP/HTTPS).

Responsibilities

  • Design and maintain scalable production infrastructure.
  • Automate operational workflows and infrastructure provisioning.
  • Enhance system reliability and performance using engineering best practices.
  • Develop CI/CD pipelines for automated application deployments.
  • Monitor and troubleshoot production systems, conducting root cause analysis.
  • Manage observability solutions including logging and alerting.
  • Collaborate with development teams for improved application reliability.

Benefits

  • Comprehensive health and wellness benefits.
  • Flexible or hybrid work arrangements.
  • Learning and certification sponsorship opportunities.
  • Access to modern cloud infrastructure and engineering tools.
  • Opportunity to work on scalable mission-critical systems.
Full Job Description
Site Reliability Engineer (SRE)

Job Title

Site Reliability Engineer (SRE)

Job Summary

We are seeking a skilled Site Reliability Engineer (SRE) to build, automate, and maintain highly available, scalable, and reliable infrastructure and applications. The ideal candidate will combine software engineering and operations expertise to improve system reliability, automate operational processes, optimize performance, and ensure service availability. You will work closely with software developers, DevOps engineers, cloud architects, and security teams to support production systems and drive operational excellence.

Key Responsibilities
  • Design, implement, and maintain highly available and scalable infrastructure for production environments.
  • Automate infrastructure provisioning, deployment, monitoring, and operational workflows.
  • Improve system reliability, availability, scalability, and performance through engineering best practices.
  • Develop and maintain CI/CD pipelines to support automated application deployments.
  • Monitor production systems, troubleshoot incidents, perform root cause analysis (RCA), and implement preventive measures.
  • Configure and manage observability solutions including logging, monitoring, tracing, and alerting.
  • Collaborate with development teams to improve application reliability and operational readiness.
  • Implement disaster recovery, backup, and business continuity strategies.
  • Optimize cloud infrastructure utilization and operational costs.
  • Develop automation scripts and tools to eliminate repetitive operational tasks.
  • Participate in incident response, on-call rotations, and post-incident reviews.
  • Document infrastructure architecture, operational procedures, and system configurations.

Required Qualifications
  • Bachelor's degree in Computer Science, Information Technology, Software Engineering, or a related field.
  • 3-6 years of experience in Site Reliability Engineering, DevOps, Cloud Engineering, or Systems Engineering.
  • Strong proficiency in Linux system administration.
  • Experience with Python, Go, Bash, or Shell scripting.
  • Hands-on experience with Docker and Kubernetes.
  • Strong understanding of networking concepts including TCP/IP, DNS, HTTP/HTTPS, and load balancing.
  • Experience with CI/CD tools such as Jenkins, GitHub Actions, GitLab CI/CD, or Azure DevOps.
  • Familiarity with cloud platforms such as AWS, Microsoft Azure, or Google Cloud.
  • Strong understanding of infrastructure automation and configuration management.

Preferred Qualifications
  • Experience with Infrastructure as Code (IaC) tools such as Terraform, Ansible, Pulumi, or CloudFormation.
  • Hands-on experience with monitoring and observability tools including Prometheus, Grafana, ELK Stack, OpenTelemetry, Datadog, or Splunk.
  • Familiarity with service mesh technologies such as Istio or Linkerd.
  • Experience managing distributed systems and microservices architectures.
  • Knowledge of security best practices, identity management, and compliance requirements.
  • Experience supporting high-availability, mission-critical production systems.

Technical Skills
  • Linux
  • Python
  • Go
  • Bash / Shell Scripting
  • Docker
  • Kubernetes
  • Terraform
  • Ansible
  • Jenkins
  • GitHub Actions
  • GitLab CI/CD
  • Prometheus
  • Grafana
  • ELK Stack
  • OpenTelemetry
  • Datadog
  • Splunk
  • Git
  • AWS / Azure / Google Cloud
  • Networking (TCP/IP, DNS, HTTP/HTTPS)
  • SQL (basic)

Soft Skills
  • Strong analytical and troubleshooting abilities.
  • Excellent problem-solving and incident management skills.
  • Effective communication and collaboration across engineering teams.
  • Ability to perform under pressure during production incidents.
  • Strong documentation and organizational skills.
  • Continuous improvement mindset and commitment to automation.

Nice to Have
  • Experience supporting AI/ML infrastructure or MLOps platforms.
  • Familiarity with container security and Kubernetes security best practices.
  • Knowledge of FinOps and cloud cost optimization.
  • Experience with chaos engineering and resilience testing.
  • Relevant certifications such as AWS Certified DevOps Engineer, Certified Kubernetes Administrator (CKA), Google Professional Cloud DevOps Engineer, or Microsoft Azure DevOps Engineer Expert.

Benefits
  • Competitive salary and performance-based incentives.
  • Comprehensive health and wellness benefits.
  • Flexible or hybrid work arrangements.
  • Learning, certification, and conference sponsorship opportunities.
  • Access to modern cloud infrastructure and engineering tools.
  • Opportunity to work on highly scalable, mission-critical systems in a collaborative and innovative environment.

Similar Jobs

More Jobs at Ova Technologies

  • Information Security Manager
    $120K — $150K *
    New York, NY 10025 (New York County)
    Information Technology
    In-Person
  • SOC Analyst
    $70K — $95K *
    Remote
    Information Technology
    Remote in New York, NY
  • SOC Analyst
    $80K — $120K *
    New York, NY 10025 (New York County)
    Information Technology
    In-Person
  • Incident Response Analyst
    $90K — $130K *
    New York, NY 10025 (New York County)
    Information Technology
    In-Person
  • Windows Administrator
    $80K — $120K *
    New York, NY 10025 (New York County)
    Information Technology
    In-Person

More Information Technology Jobs

Find similar Site Reliability Engineer (SRE) jobs: