OverviewWe are designing the grid of the future!
We are seeking an experienced Senior Site Reliability Engineer (SRE) to join our engineering organization. The ideal candidate will combine strong software engineering and cloud infrastructure expertise to improve the reliability, scalability, security, and operational efficiency of our systems. This role will focus on building resilient AWS and Kubernetes platforms, defining and measuring service reliability, automating operational work, and improving how we detect, respond to, and learn from production incidents.
The Senior SRE will work closely with software engineering, platform, security, QA, and product teams to establish reliability standards and ensure production services can scale safely as the business grows.
Responsibilities
Howyou can make an impact:
- Service Reliability & Availability: Define, measure, and improve service reliability using service-level indicators (SLIs), service-level objectives (SLOs), error budgets, availability targets, and capacity planning.
- Observability: Build and maintain monitoring, logging, tracing, dashboards, and alerting using tools such as CloudWatch, Prometheus, Grafana, New Relic, or similar platforms. Ensure alerts are actionable and aligned to customer and service impact.
- Incident Response: Participate in and improve production incident response, including on-call practices, troubleshooting, escalation, communication, root-cause analysis, and blameless post-incident reviews.
- Operational Excellence: Improve runbooks, documentation, production readiness reviews, change management, operational standards, and engineering practices that reduce risk and improve system maintainability.
- Cost & Efficiency: Optimize infrastructure for reliability and performance while maintaining responsible cloud spend and supporting FinOps initiatives.
- Performance & Capacity: Analyze system performance, resource utilization, latency, throughput, and growth trends; identify bottlenecks and implement scalable solutions before they become production issues.
- Automation & Toil Reduction: Identify repetitive operational work and replace it with reliable automation using Python, Bash, CI/CD tooling, and platform APIs.
- CI/CD & Release Reliability: Build and improve deployment pipelines that support safe, repeatable releases through automated testing, validation, progressive delivery, rollback strategies, and deployment observability.
- Security & Compliance: Apply cloud security and operational controls including least-privilege IAM, encryption, network security, secrets management, patching, auditability, and compliance requirements.
- Infrastructure as Code: Design, develop, review, and maintain reusable Terraform modules and infrastructure-as-code patterns for AWS environments.
- Kubernetes Reliability: Operate and improve Kubernetes-based platforms, including EKS clusters, workloads, Helm deployments, autoscaling, upgrades, resource management, and workload resilience.
- Cloud Platform Engineering: Architect, operate, and optimize AWS services such as EC2, S3, RDS, EKS, Lambda, VPC, IAM, Route 53, and related services.
- Resilience & Disaster Recovery: Design and validate fault-tolerant architectures, backup strategies, recovery procedures, and disaster recovery capabilities. Conduct reliability testing and failure exercises where appropriate.
- Cross-Functional Collaboration: Partner with development teams to improve application operability, reliability, instrumentation, deployment patterns, and production readiness.
Qualifications
Bring your passion, heres whats needed:
Required Skills and Qualifications
- Experience: 5+ years of experience in Site Reliability Engineering, DevOps, platform engineering, cloud infrastructure, or a closely related discipline, including significant production ownership.
- AWS: Strong hands-on experience designing and operating production workloads in AWS, including networking, IAM, compute, storage, databases, DNS, and managed Kubernetes.
- Terraform: Advanced proficiency with Terraform, including reusable modules, remote state, dependency management, environment design, code review, and infrastructure lifecycle management.
- Observability: Experience with metrics, logs, traces, dashboards, alerting, and production telemetry using platforms such as Prometheus, Grafana, CloudWatch, New Relic, ELK/OpenSearch, or similar tools.
- Reliability Engineering: Practical knowledge of SRE concepts such as SLIs, SLOs, error budgets, capacity planning, fault tolerance, graceful degradation, and reducing operational toil.
- Kubernetes: Deep understanding of Kubernetes architecture and operations, including EKS, Helm, workload scheduling, networking, storage, autoscaling, upgrades, and troubleshooting.
- Linux & Systems: Strong Linux systems knowledge and the ability to diagnose issues involving CPU, memory, disk, networking, processes, DNS, and application dependencies.
- Programming & Automation: Proficiency in Python, Bash, Go, or another general-purpose language used to build operational tooling and automation.
- Incident Management: Experience troubleshooting complex production incidents and contributing to incident response, root-cause analysis, postmortems, and corrective-action tracking.
- CI/CD: Experience designing or operating CI/CD systems such as GitHub Actions, Jenkins, GitLab CI, Argo CD, or comparable tooling.
- Security: Working knowledge of cloud security best practices, including IAM, encryption, secrets management, network segmentation, vulnerability management, and audit controls.
- Communication: Strong written and verbal communication skills, with the ability to collaborate effectively across engineering and business teams.
Preferred Qualifications
- AWS certifications such as AWS Certified Solutions Architect - Professional or AWS Certified DevOps Engineer - Professional.
- Experience with GitOps practices and tools such as Argo CD or Flux.
- Experience designing or participating in formal on-call rotations and incident management programs.
- Familiarity with chaos engineering, resilience testing, or game-day exercises.
- Experience with service meshes, distributed systems, and microservice architectures.
- Knowledge of database operations for technologies such as Amazon RDS, DynamoDB, and PostgreSQL.
- Experience with multi-account AWS environments, landing zones, governance, or large-scale cloud platform design.
- Experience with infrastructure cost optimization, cloud financial management, or FinOps practices.
- Familiarity with security and compliance frameworks such as SOC 2, ISO 27001, PCI DSS, or similar standards.
What Success Looks Like
- Production services become more measurable, reliable, and resilient over time.
- Operational toil and recurring incidents are reduced through engineering and automation.
- Teams have clear SLOs, useful dashboards, actionable alerts, and well-understood operational ownership.
- Infrastructure and deployment changes are repeatable, observable, secure, and low risk.
- Incidents produce meaningful learning and durable improvements rather than recurring fixes.
- Cloud capacity and cost are proactively managed without compromising reliability.
Why Join Us?
- Work on modern cloud and reliability engineering challenges in a fast-paced, innovative environment.
- Help shape reliability standards and engineering practices for mission-critical systems.
- Collaborate with talented engineers across software, cloud, security, and product disciplines.
- Competitive salary, comprehensive benefits, and opportunities for professional growth.
- Flexible remote or hybrid work options.
Be a part of an innovative team shaping the grid of the future through advanced energy intelligence. For more than half a century, Electric Power Engineers (EPE) has partnered with power and energy clients across the globe, providing consulting expertise and energy intelligence software solutions for complex engineering and grid modeling challenges. As leaders in the renewables space, we are focused on building a modern, secure, and resilient grid. Join us in making an impact on the communities we serve and the environment in which we live. Together we can transform the future of energy.
How we support you:
- Comprehensive health and wellness benefits including medical, dental, and vision with 100% premium coverage foryou
- Generous PTO and paid holidays
- MyShare Employee Ownership Program
- Work with industry leaders
- 401K, up to a 4% match (100% vested from day 1)
Location: This position will be located in City, State
Travel: Occasional travel may be needed (10% or less)