Role SummaryCloud Reliability operates as a center of excellence for cloud and reliability engineering, delivering secure, scalable, and resilient cloud platforms to teams across the firm. The Principal Cloud Reliability Engineer is a senior individual contributor responsible for defining, building, and operating enterprise cloud foundations, with a strong emphasis on reliability, observability, automation, and operational excellence. The role combines enterprise technical authority with hands-on execution, ensuring that cloud platforms, particularly AWS Landing Zone, enable teams to deliver end-to-end applications that meet the firm's availability, resilience, security, and risk expectations. The Principal serves as a design authority, site reliability engineering leader, and escalation point for complex cloud platform challenges.
ResponsibilitiesThe Principal Cloud Engineer (Network/Security) owns the architecture, implementation, and reliability outcomes of enterprise cloud platforms. The role balances strategic leadership with direct contribution, driving reliability standards while remaining hands-on in critical areas of platform engineering, automation, observability, and incident response.
- Lead the design, architecture, and evolution of enterprise cloud platforms, including AWS Landing Zone.
- Ensure cloud foundations are designed for high availability, fault tolerance, recoverability, security, and operational resilience.
- Design and implement cloud solutions using AWS services such as Amazon EC2, Amazon S3, Amazon ECS, Amazon EKS, Elastic Load Balancing, Amazon RDS, Route 53, AWS Lambda, and Amazon API Gateway.
- Design and automate advanced AWS networking solutions, including virtual private clouds, Transit Gateway, VPC peering, AWS PrivateLink, and AWS Direct Connect.
- Build and maintain infrastructure as code using Terraform, CloudFormation, Ansible, and Git-based workflows.
- Own site reliability engineering outcomes across cloud platforms and drive adoption of reliability engineering practices.
- Standardize reusable modules, reference architectures, and platform patterns that promote consistent and reliable deployments at scale.
- Define and enforce reliability standards, including availability targets, recovery expectations, resilience patterns, and operational readiness criteria.
- Ensure instrumentation, monitoring, logging, alerting, and service health measures are embedded into platforms and services.
- Act as the escalation point for complex incidents, leading root-cause analysis and long-term remediation.
- Design and implement guardrails that enable secure, reliable, and self-service cloud adoption.
- Enable application teams to own end-to-end services while meeting platform reliability and operational standards.
- Participate in an Agile delivery model, contributing to stories and epics, tracking work, and supporting sprint releases.
- Partner closely with Application Development, Development Services, Enterprise Architecture, Enterprise Security, and infrastructure teams.
- Mentor engineers and promote a culture of operational excellence, continuous improvement, and technology modernization.
Business Knowledge- Deep understanding of how reliability, availability, recoverability, and platform performance affect business outcomes and client experience.
- Ability to balance delivery speed with security, resilience, risk, cost, and operational sustainability.
- Experience operating in regulated, risk-aware environments with strong security and compliance requirements.
- Ability to translate enterprise strategy into practical cloud platform standards, roadmaps, and engineering priorities.
- Makes decisions aligned with enterprise technology strategy while improving incident recovery, reducing repeat issues, and strengthening platform stability.
Qualifications
Required:- Bachelor's degree or the equivalent combination of education and relevant experience AND 10+ years of experience designing and operating cloud infrastructure with senior-level impact.
- Deep hands-on experience with AWS, with working knowledge of Microsoft Azure.
- Expert knowledge of cloud infrastructure, container platforms, serverless deployments, operating systems, networking, and identity and access management.
- Expert knowledge of reliability engineering concepts, including availability, resilience, observability, recovery, incident response, and operational readiness.
- Strong experience with AWS Landing Zone or comparable enterprise cloud foundation capabilities.
- Advanced infrastructure automation experience using Terraform, CloudFormation, Ansible, scripting, APIs, continuous integration and delivery pipelines, and secure development and operations practices.
- Proven ability to design, build, and operate enterprise-scale, highly reliable cloud platforms.
- Strong troubleshooting skills across infrastructure, networking, platform services, automation, and reliability layers.
- Strong written and verbal communication skills, with the ability to serve as a design authority and influence teams without formal line management.
- Experience mentoring engineers and leading technical decisions across organizational boundaries.
Preferred:- Cloud or site reliability engineering certifications
- Experience in financial services - asset management, investment banking, or fintech.
FINRA RequirementsFINRA licenses are not required and will not be supported for this role.
Work FlexibilityThis role is eligible for hybrid work, with up to three days per week from home.
Applicants for employment in the US must have work authorization that does not now or in the future require sponsorship of a visa for employment authorization in the United States (e.g., H1-B visa, F-1 visa (OPT), TN visa or any other non-immigrant work status).Base Salary RangesPlease review the job posting for the location of this specific opportunity.
$159,000.00 - $272,000.00 for the location of: Maryland, Colorado, Washington and remote workers
$175,000.00 - $299,000.00 for the location of: Washington, D.C.
$199,000.00 - $339,000.00 for the location of: New York, California
Placement within the range provided above is based on the individual's relevant experience and skills for the role. Base salary is only one component of our total compensation package. Employees may be eligible for a discretionary bonus, which is determined upon company and individual performance.
BenefitsWe value your goals and needs, at work and in life. As an associate, you'll be supported with resources, benefits, and work-life balance so you can thrive in ways that matter to you.
Featured employee benefits to enrich your life:
- A generous retirement plan
- Health and wellness benefits, including online therapy
- Paid time off for vacation, illness, medical appointments, and volunteering days
- Family care resources, including fertility and adoption benefits
Learn more about our benefits.