Location Designation: Hybrid - 3 days per quarter
Role Overview:
New York Life is seeking an experienced Site Reliability Engineer (SRE) to join the Cloud Platform Engineering Team and provide both technical leadership and day-to-day management for multi cloud solution delivery (AWS, GCP, Azure). This role will partner with application teams, architects, security, operations, and business stakeholders to design, deliver, and continuously improve secure, scalable, reliable, and cost-effective solutions on AWS.
The role will build new IaC artifacts and governed curated registries which will be consumed by NYL application teams. The goal is to shift from non-standard artifacts to standard platform patterns offering a full stack deployment of infrastructure which requires little to zero intervention by application development teams. Most of our applications are deployed on AWS and we leverage GCP for building agentic AI agents. Azure supports our Digital Workplace Services.
The ideal candidate combines strong hands-on cloud engineering experience with the ability to guide technical direction, mentor engineers, build trusted stakeholder relationships, and drive execution across multiple workstreams.
Successful reliability outcomes are likely to implement and extend on DevOps and Agile ways of working and associated automation approaches. These are underpinned by the site reliability engineer's solid understanding of systems, production environments, operational insights, incident management, on-premises, cloud and hybrid world. The nature of the work involved means that the site reliability engineer will directly engage with customer teams but will also work on reliability initiatives that span multiple teams.
The SRE collaborates closely with product owners and teams, architects, IT service management, software developers, security and network engineers, as well as other subject matter experts and roles, particularly in infrastructure and operations. Being an approachable team player and a good communicator is therefore crucial for success, and a willingness to lead initiatives is important.
What You'll Do:
• Support our Cloud Business Office by facilitating Well Architected reviews for new cloud deployments or cloud migration candidates.
• Lead the design and delivery of repeatable, reliable cloud solutions using Terraform, GitHub, and related automation tools
• Develop and deploy full stack patters to support ongoing application development, data platform capabilities, and AI services (RAG & Agentic)
• Partner with application and product teams to assess requirements, recommend AWS architecture patterns, and support successful cloud implementations
• Support the onboarding and lifecycle management of cloud services by producing standard IaC/Terraform Modules.
• Define and mature SRE practices, including SLO/SLI frameworks and error-budget governance.
• Design and implement automation solutions using Java, JavaScript, APIs, SQL, and Terraform.
• Deliver application-level fixes and enhancements through disciplined software engineering.
• Focus on key reliability and performance indicators: uptime, system throughput, system output, and download rate/application load speed.
• Lead the shift from non-standard application platforms to standard software artifacts (Terraform modules, secure base images, YAML templates, Java libraries) integrated into CI/CD pipelines, creating reusable patterns and reducing repetitive configuration and coding tasks.
• Review cloud platform designs, operational issues, and delivery risks; provide practical recommendations and follow-through to resolution
• Advise teams on cloud cost optimization, resource utilization, governance, and platform standards
• Documentation: Maintain detailed records of issues, actions taken, and outcomes to support continuous improvement efforts.
• Collaboration: Work closely with other IT teams and external vendors to resolve complex issues and implement solutions.
• Process Improvement: Identify opportunities to improve support processes and implement best practices to enhance overall efficiency.
• Training: Provide training and guidance to IT staff on network and platform support techniques and best practices.
What You'll Bring:
• 8+ years of overall IT, infrastructure, platform engineering, or cloud engineering experience
• 5+ years of experience with AWS cloud platforms and services
• 3+ years of experience designing, automating, and operating scalable, secure, and highly available cloud solutions
• Ability to produce production grade IaC and perform code reviews.
• Demonstrated ability to lead, mentor, and coordinate engineers or technical contributors in a matrixed environment
• Strong business and technical acumen with the ability to influence decisions, manage tradeoffs, and communicate with leadership
• Strong facilitation, planning, and execution skills for cross-functional workshops, technical reviews, and delivery initiatives
• Solid understanding of SDLC, DevOps, change management, incident response, and production operations practices
• Ability to work directly with business stakeholders, architects, security teams, product owners, and engineering teams to translate needs into practical cloud solutions
• AWS Solutions Architect, AWS SysOps Administrator, AWS DevOps Engineer, or equivalent AWS certification preferred
Desired Skills:
• Experience with cloud governance, landing zones, identity and access management, networking, compute, storage, and security control patterns
• Hands-on experience with infrastructure as code practices using Terraform, CloudFormation, AWS CDK, or similar AWS-focused tools
• Experience with Kubernetes, Amazon EKS, container platforms, CI/CD pipelines, and observability tools is a plus
• Practical understanding of infrastructure technologies, including compute, network, storage, DNS, load balancing, encryption, logging, and monitoring
• Ability to relate cloud platform capabilities to software engineering, application modernization, resiliency, and developer experience needs
• Working knowledge of DevOps and delivery tools such as GitHub, GitHub Actions, Jfrog Artifactory, Harness, Teraform, or similar platforms
• Practical knowledge of scripting or programming languages such as Python, PowerShell, .NET, C#, Java, or similar
• Strong time management, prioritization, coaching, and problem-solving skills with the ability to balance delivery commitments and operational needs
• Background with cloud cost management, FinOps practices, compliance requirements, and risk-based decision making is a plus
Pay Transparency
Salary Range: $147,500-$211,000
Overtime eligible: Exempt
Discretionary bonus eligible: Yes
Sales bonus eligible: No
Actual base salary will be determined based on several factors but not limited to individual's experience, skills, qualifications, and job location. Additionally, employees are eligible for an annual discretionary bonus. In addition to base salary, employees may also be eligible to participate in an incentive program.
Our Benefits
We provide a full package of benefits for employees - and have unique offerings for a modern workforce, including leave programs, adoption assistance, and student loan repayment programs. Based on feedback from our employees, we continue to refine and add benefits to our offering, so that you can flourish both inside and outside of work.Click hereto discover more about our comprehensive benefit options or visit our NYL Benefits Site.
Job Requisition ID: 94732