Site Reliability Engineer

Schonfeld

$175K — $225K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years of cloud automation or site reliability engineering experience
  • Proficient in Python and software development best practices
  • Experience with REST APIs and asynchronous architectures
  • Familiar with relational (Postgres, MySQL) and NoSQL (DynamoDB, Elastic Search) databases
  • Knowledgeable in AWS services, Kubernetes, and modern CI/CD practices
  • Strong problem-solving skills and analytical thinking
  • Excellent communicator, able to explain technical concepts to non-technical audiences

Responsibilities

  • Define reliability standards, SLOs, and incident response protocols for AI operations
  • Manage observability and reliability of AI platforms
  • Ensure high availability and efficiency of AI processes and systems
  • Provide technical support and root cause analysis for users
  • Contribute to code quality and improvement in development practices

Benefits

  • Competitive benefits package including performance bonuses
  • Opportunities for hands-on work in cutting-edge AI technology
  • A dynamic and fast-paced work environment with evolving product requirements
Full Job Description
The Role

We're looking for a Site Reliability Engineer to join our AI Technology team and play a critical role in ensuring the stability and scalability of the firm's internal Agentic AI platform. This is a high-impact, hands-on engineering position where you'll contribute code, automation workflows, observability, metrics, and support users as issues arise. You'll work in a dynamic, fast-paced environment where product requirements evolve, and business stakeholders are deeply engaged.

What you'll do
  • You will set the reliability standards for Enterprise AI, defining Service Level Objectives (SLOs), error budgets, and custom incident response runbooks.
  • You will own the observability, incident response, reliability, and scalability of our AI platform.
  • Ensure that our agents, gateways, LLM proxies, and RAG pipelines operate with high availability, accuracy, and financial efficiency.
  • Support users in a dedicated help channel when issues arise - investigating root causes, providing solutions, monitoring the status of upstream dependencies, and communicating updates back to users.
  • You'll also be involved in the team's code quality and best practices, identify gaps in development lifecycles, and continuously improve both the core platform and how developers build with it.

What you need:
  • 5+ years of professional cloud automation or site reliability engineering experience, or similar
  • Familiarity with software development best practices and experience creating applications in Python
  • Experience working creating and integrating REST APIs, and event-driven and asynchronous architectures
  • Experience working with both relational (Postgres, MySQL) and NoSQL databases (DynamoDB, Elastic Search)
  • Experience with AWS cloud services (S3, OpenSearch, DynamoDB, etc.), Kubernetes-driven deployments, Github Actions, and modern CI/CD pipelines
  • A passion for AI, creative ideas for how it can be a tool for problem solving and productivity
  • A thoughtful approach to problem-solving; ability to break down complex issues methodically
  • Excellent communication skills; Able to translate technical concepts for non-technical stakeholders

We'd love if you had:
  • Experience developing agentic AI applications, harnesses, MCP servers, and workflows
  • Familiarity with or experience building AI agents, Retrieval-Augmented Generation (RAG) pipelines, or agentic tools
  • Experience with observability tooling such as DataDog

The base pay for this role is expected to be between $175,000 and $225,000. The expected base pay range is based on information at the time this post was generated. This role may also be eligible for other forms of compensation such as a performance bonus and a competitive benefits package. Actual compensation for the successful candidate will be determined based on a variety of factors such as skills, qualifications, and experience.

The base pay for this role is expected to be between $150,000 and $180,000. The expected base pay range is based on information at the time this post was generated. This role may also be eligible for other forms of compensation such as a performance bonus and a competitive benefits package. Actual compensation for the successful candidate will be determined based on a variety of factors such as skills, qualifications, and experience.

Similar Jobs

More Jobs at Schonfeld

More Information Technology Jobs

Find similar Site Reliability Engineer jobs: