Job DescriptionOverview of Job Function:As an Associate SRE, you will be a hands-on contributor on the SRE team, supporting reliability, automation, observability, and incident response efforts across the organization. Working under the guidance of senior and principal SREs, you will help implement SLIs/SLOs, reduce manual toil through automation, improve telemetry, and participate in on-call rotations and blameless postmortems. You will be expected to bring a growth mindset, a bias toward "everything-as-code," and a willingness to ask questions, dig into unfamiliar systems, and continuously improve. This role offers strong mentorship and exposure to cross-organizational collaboration in the true spirit of DevOps.
Principal Duties and Essential Responsibilities:- Implement and maintain SLIs and SLOs for assigned services, working alongside senior SREs
- Write automation scripts and pipelines to reduce manual toil within CloudOps
- Contribute to Infrastructure and Configuration as Code (e.g. Terraform, Helm, Ansible)
- Build, tune, and maintain dashboards, alerts, and telemetry in observability tools
- Partner with R&D and Operations teams to enhance telemetry and improve service reliability
- Participate in the on-call rotation and respond to incidents under the guidance of senior engineers
- Support the design and implementation of automated incident response workflows and runbooks
- Participate in blameless postmortems and follow up on assigned engineering action items
- Shadow senior SREs during complex outages, architectural reviews, and fire drills
- Build AI-powered solutions to operational problems, with a particular focus on improving availability (e.g. agentic incident triage, predictive alerting, automated remediation, AI-assisted runbooks)
- Continuously learn cloud, Kubernetes, and SRE best practices and apply them to daily work
- Take ownership of assigned tasks and communicate progress, blockers, and learnings clearly
Minimum Requirements:- B.S. degree in Computer Science, Engineering, or related field (or equivalent practical experience, including bootcamps or demonstrable self-taught work)
- 1-3 years of professional experience in software engineering, SRE, DevOps, or systems administration
- Exposure to at least one public cloud (AWS or Azure)
- Exposure to containers and Kubernetes fundamentals (Docker, Helm, YAML)
- Familiarity with Linux and shell scripting; exposure to Windows/PowerShell a plus
- Comfortable using Git for branching, pull requests, and merges
- Strong troubleshooting curiosity and willingness to dig into unfamiliar systems
- Willingness to participate in an on-call rotation
- Strong written and verbal communication skills
- Growth mindset - actively seeks feedback and continuously improves
- High level of personal ownership, accountability, and drive
Preferred Skills:- Entry-level cloud certification (AWS Cloud Practitioner, Developer Associate, or SysOps Associate; Azure Fundamentals)
- Familiarity with at least one observability tool (Datadog, Prometheus/Grafana, Splunk, SumoLogic, AppDynamics, etc.)
- Working proficiency in at least one scripting or programming language (Python, Bash, Go, Java, C#, etc.)
- Exposure to Infrastructure as Code (Terraform, CloudFormation, Ansible, Pulumi)
- Exposure to CI/CD pipelines (Harness, GitHub Actions, Jenkins, GitLab CI, etc.)
- Internship, co-op, or project experience with SaaS, PaaS, or multi-tenant environments
- Familiarity with incident management tools (OpsGenie, PagerDuty, ServiceNow)
- Exposure to Agile and DevOps practices
- Academic, open-source, or personal projects demonstrating automation, reliability, or observability work
- Experience with AI agentic coding tools (e.g. Claude Code, Cursor, GitHub Copilot) and integrating AI capabilities into solutions and workflows
- Detail oriented and organized with the ability to manage multiple priorities
- Positive, collaborative energy and a desire to contribute to a growing team
#LI-KD1
About the Team2025 Benefits Offering