Horizon3.ai

Staff Site Reliability Engineer

Horizon3.ai • $199K — $270K *
US-AnywhereRemote in United States
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience in site reliability engineering or related fields
  • Proven track record in designing large-scale distributed systems
  • Strong expertise in reliability engineering, observability, and incident management
  • Hands-on experience establishing service level indicators (SLIs) and service level objectives (SLOs)
  • Backend development experience with Python and Terraform, or similar tools
  • Excellent communication skills for technical documentation and incident response

Responsibilities

  • Own and develop the engineering-wide SRE strategy and standards
  • Lead cross-functional teams to enhance reliability and operational readiness
  • Establish meaningful diagnostics for service ownership, SLIs, and error budgets
  • Design observability standards for effective service health monitoring
  • Create templates for dashboards and streamline incident management processes
  • Manage end-to-end reliability initiatives across different functions
  • Shape the operational model and growth of the SRE team

Benefits

  • Inclusive team culture promoting diversity
  • Opportunities for career growth and professional development
  • Collaborative and innovative work environment
  • Flexible hybrid and remote work arrangements
  • Comprehensive health, vision, and dental insurance benefits
  • Generous parental leave and flexible vacation policy
Full Job Description
We are seeking a hands-on Staff Site Reliability Engineer to own and evolve the reliability strategy, operating model, and engineering-wide standards supporting our platform. This is a foundational role for an experienced engineer who will set technical direction across teams, lead the highest-impact reliability initiatives, and establish the practices and systems that enable engineering to operate production services safely.

What You'll Do
  • Own and evolve the engineering-wide SRE strategy, operating model, and reliability standards, aligning them to customer impact, business priorities, and risk.
  • Lead cross-functional alignment across Infrastructure, product, service, security, and business stakeholders to improve reliability, observability, incident response, and operational readiness across multiple teams.
  • Establish an organization-wide approach to service ownership, meaningful SLIs and SLOs, and error budgets for critical customer paths and services.
  • Define and drive adoption of observability standards across pipelines and platform components that report on service health, performance and operational risk.
  • Set the standard for dashboards, actionable alerting, runbooks, and escalation paths.
  • Drive end to end complex cross-functional reliability initiatives.
  • Set and raise the engineering wide bar for incident management, incident command, on-call health, post-incident learning, and recovery readiness.
  • Shape the technical direction, operationable model, and growth path of the SRE function.
  • Participate in a 24/7 on-call rotation and help design an on-call model that is sustainable, appropriately staffed, and continuously improved.


What You'll Bring
  • Experience designing, operating, and troubleshooting large scale distributed systems in production environments.
  • Deep knowledge of reliability engineering, observability, incident management, and production operations, with demonstrated ability to turn that knowledge into standards and practices adopted by others
  • Experience in establishing SLIs, SLOs, actionable alerts, observability, and service ownership.
  • Backend experience building backend systems and automation that reduce optional toil, strengthen safeguards, and operational efficiency.
  • Experience in leading high severity incidents and improving incident response programs.
  • Excellent written and verbal communication skills including technical designs, runbooks, postmortems, and operational documentation.


Required Tech Stack Experience
  • Python and Terraform (Infrastructure as Code), or equivalent automation and infrastructure-as-code tools.
  • Experience with Observability tools such as Datadog, New Relic, Grafana, or equivalent platforms.
  • Experience operating production services in AWS and Kubernetes
  • Experience with CI/CD pipelines such as Gitlab CI, ArgoCD, or GitOps workflows..


Travel Required

We are a fully remote company, and this job may require up to 10% of travel to be successful. Travel primarily consists of team off-sites and in-person project kick-offs.

Perks of Horizon3
  • Inclusive Team: We value diversity and promote an inclusive culture where everyone can thrive.
  • Growth Opportunities: Be part of a dynamic and growing team with numerous career development opportunities.
  • Innovative Culture: Work in a collaborative environment that encourages creativity and out-of-the-box thinking.
  • Hybrid & Remote Work: We embrace a mix of remote and hybrid work models depending on role and location, including our Chicago office, where some roles require regular in-office presence.
  • Competitive Compensation: We offer competitive salary, equity and benefits. Our benefits include health, vision & dental insurance for you and your family, a flexible vacation policy, and generous parental leave.


Compensation and Values

At Horizon3, we believe that our people are our greatest asset, and our compensation philosophy reflects this core value. We are committed to fostering an environment where all employees feel valued, respected, and rewarded for their contributions. Our compensation structure is designed to be fair, competitive, and transparent, ensuring that every team member is recognized and compensated equitably across roles, levels, and locations.

In accordance with various State's transparency regulations, we provide the following salary range information for this position:
  • Base salary range: $199,750 - $270,000 annually. The exact salary will be determined based on the selected candidate's location, qualifications, experience, and relevant skills.
  • Additional compensation: All full-time roles are eligible for an equity package in the form of stock options.


Other Duties

Please note this job description is not designed to cover or contain a comprehensive listing of activities, duties or responsibilities that are required of the employee. Duties, responsibilities, and activities may change at any time with or without notice.

Application Note

In any materials you submit, you may redact or remove age-identifying information such as age, date of birth, or dates of school attendance or graduation. You will not be penalized for redacting or removing this information.

About Horizon3.ai

Horizon3.ai is a California-based software company that provides artificial intelligence (AI) solutions for businesses. The company was founded in 2018 and is headquartered in Long Beach, California. Horizon3.ai offers a range of AI-powered products, including chatbots, virtual assistants, and predictive analytics tools. The company's solutions are designed to help businesses automate their operations, improve customer engagement, and gain insights from their data. Horizon3.ai serves clients in a variety of industries, including healthcare, finance, and retail.
Learn more about Horizon3.ai
Size
50 employees
Industry
Net Income
-$100,000
Founded
2018
5 Year Trend
+50%
Revenue
$500,000

Similar Jobs

More Jobs at Horizon3.ai

More Information Technology Jobs

Find similar Staff Site Reliability Engineer jobs: