What you'll do:As a Site Reliability Engineer at Zefr, you'll join our Engineering team and grow your skills in cloud infrastructure, CI/CD, Observability, and core SRE concepts, helping deliver high-quality, reliable, and scalable solutions. You'll work closely with the rest of Zefr's Engineering and Data Science teams, ensuring the infrastructure behind our services is robust, efficient, and scalable.
We are looking for a passionate engineer to come collaborate and grow alongside our experienced SRE team. You bring curiosity, a strong work ethic, and a passion for continuous improvement and automation.
- Support and build systems and tools that enable other engineers to deploy and manage product features both quickly and safely.
- Help deploy and support our multi-cloud, micro-service architecture, deployed via Github Actions, ArgoCD & Kubernetes.
- Write and maintain Infrastructure as Code with Terraform and Terragrunt, contributing changes through pull requests and code review.
- Build and improve CI/CD pipelines and release workflows.
- Help maintain the health of production environments, including monitoring application performance and resource utilization.
- Participate in 24/7 on-call rotation, responding to system performance issues and outages alongside senior teammates.
- Debug issues at the application and infrastructure level.
- Write and maintain clear documentation and runbooks.
- Contribute to our DevOps culture and philosophy of continuous improvement.
Technology Stack at Zefr:- Cloud Providers: Google Cloud Platform (primary), Amazon Web Services
- Infrastructure as Code (IaC): Terraform, Terragrunt
- Containerization & Orchestration: Docker, Kubernetes (GKE)
- CI/CD: GitHub Actions, Argo CD
- Observability: Prometheus, OpenTelemetry, Chronosphere, Pagerduty
- Application Languages/Frameworks: Python, FastAPI, Flask, Node.js, React
- Workflow Orchestration: Apache Airflow, Ray
- Relational Databases: PostgreSQL
- NoSQL Databases: DynamoDB
- Search Databases: OpenSearch
- Data Warehousing: Snowflake
What we're looking for:- 1-3 years of experience supporting Cloud Infrastructure in a production environment using AWS and/or GCP
- Hands-on experience with containers and Kubernetes
- Competency in Python and shell scripting
- Strong problem-solving skills
- Strong written and verbal communication, organization, and documentation skills
Nice to have:- Familiarity with modern CI/CD pipelines and GitOps (Github Actions, GitLab, Argo CD)
- Exposure to Monitoring and Observability tools (Prometheus, Grafana, Chronosphere, Datadog, OpenTelemetry)
- Experience collaborating across departments to execute complex, high-impact projects.
Benefits (for US based employees):- Flexible PTO
- Medical, dental, and vision insurance with FSA options
- Company-paid life insurance
- Paid parental leave
- 401(k) with company match
- Professional development opportunities
- 13 paid holidays off
- Summer Fridays (we leave early)
- In-office and hybrid work options available
- In-office lunches and lots of free food
- Optional in-person and virtual events (we like to celebrate!)
Compensation (for US based employees):The anticipated salary for this position is between $100,000 and $130,000. Within the range, individual pay is determined by factors such as job-related skills, experience, and relevant education or training. If your compensation expectations fall outside of this range, it may still be worth having a conversation.