Full Job Description
Job Summary The Lead Cloud Solution Architect will serve as the senior technical authority for a complex, regulated life sciences cloud environment undergoing migration from Azure to AWS. The role will provide hands-on architectural leadership across AWS and GCP cloud operations, engineering, modernization, automation, security, reliability, observability, and continuous improvement. The environment is predominantly AWS, with a meaningful GCP footprint and a smaller Azure presence. The successful candidate will work directly with Cloud Operations leadership and delivery teams and will own technical direction across the engagement.
Key Responsibilities • Serve as the primary technical authority and Lead Cloud Architect for the engagement. • Own architecture and technical direction across a large AWS multi-account environment. • Provide hands-on expertise with AWS Control Tower, AWS Organizations, Service Control Policies (SCPs), IAM, and IAM Identity Center. • Architect and support AWS compute services including EC2, Lambda, ECS, EKS, and Fargate. • Design and support database and storage architectures using RDS, DynamoDB, S3, EBS, EFS, and FSx. • Design and manage AWS networking and hybrid connectivity using Transit Gateway, Direct Connect, VPN, Route 53, and related technologies. • Implement and support AWS security services including Security Hub, GuardDuty, Inspector, and AWS Config. • Design backup, disaster recovery, and multi-region resiliency strategies using AWS Backup and related services. • Lead AWS Well-Architected reviews and drive remediation of identified findings. • Design secure, resilient, scalable, and cost-optimized cloud architectures. • Provide architectural oversight for cloud migrations, modernization, containerization, and serverless adoption. • Support AWS Bedrock and SageMaker infrastructure and operational requirements. • Provide architectural and operational guidance for GCP environments, including networking, IAM, compute, storage, security, monitoring, and cost management. • Support cross-cloud networking, governance, security, and observability between AWS and GCP. • Work directly with GCP workloads and projects and help establish consistent engineering and governance standards across AWS and GCP. • Drive Terraform adoption and migration of manually configured resources into version-controlled Infrastructure as Code. • Establish reusable Terraform modules, reference architectures, and standardized cloud patterns. • Provide technical leadership for GitLab CI/CD pipelines. • Develop operational automation using Python, boto3, Lambda, EventBridge, and Step Functions. • Identify repetitive operational activities and develop automation to reduce manual effort and operational toil. • Establish and mature SRE practices, including SLOs, SLIs, error budgets, and reliability engineering. • Provide architectural leadership for Datadog observability across AWS and GCP. • Guide dashboard design, alert tuning, log aggregation, APM, and automated remediation. • Develop standardized platform engineering patterns and self-service golden paths. • Partner with L1-L3 Operations teams to resolve complex incidents and recurring architectural issues. • Lead or support root-cause analysis for major incidents. • Design cloud environments using least-privilege and security-by-design principles. • Provide technical leadership for IAM, secrets management, encryption, vulnerability remediation, and cloud security posture. • Support cloud infrastructure operating within regulated environments and partner with compliance and security teams to maintain audit-ready infrastructure and documentation. • Ensure infrastructure changes follow appropriate change-control and validation processes. • Identify opportunities to improve cloud cost efficiency through rightsizing, storage optimization, tagging, cost allocation, and commitment optimization. • Balance performance, resilience, security, and cost when making architectural recommendations. • Develop and maintain a technical improvement roadmap for the cloud environment. • Act as the primary technical counterpart to Cloud Operations leadership. • Lead architectural discussions and present recommendations to technical and executive stakeholders. • Provide technical direction to L1, L2, and L3 cloud operations and engineering resources. • Coordinate specialist expertise when required while retaining ownership of the overall technical solution. • Participate in operational reviews, architecture discussions, and strategic roadmap planning. • Translate business and compliance requirements into executable cloud architecture and engineering initiatives.
Required Qualifications • Expert-level AWS architecture experience equivalent to a senior or principal-level AWS Cloud Architect. • Extensive hands-on experience with complex enterprise AWS environments. • Strong experience with AWS Control Tower and multi-account architectures. • Strong knowledge of AWS Organizations and Service Control Policies (SCPs). • Strong hands-on experience with AWS networking and hybrid connectivity. • Strong expertise in IAM, cloud security, and security-by-design principles. • Extensive experience with containers and serverless architectures. • Strong experience with AWS storage and database architecture. • Experience designing and implementing disaster recovery and business continuity solutions. • Strong knowledge of the AWS Well-Architected Framework. • Experience leading cloud migration and modernization initiatives. • Strong experience with cloud cost optimization and FinOps practices. • Strong hands-on experience with Terraform and cloud automation. • Experience with enterprise monitoring and observability solutions. • Strong working expertise in GCP architecture, networking, IAM, compute, storage, security, monitoring, and cost management. • Experience supporting cross-cloud AWS and GCP environments. • Experience with Python and cloud automation using boto3, Lambda, EventBridge, or Step Functions. • Experience with GitLab CI/CD and Infrastructure as Code practices. • Experience establishing SRE practices, including SLOs, SLIs, error budgets, and reliability engineering. • Experience with Datadog observability, including dashboards, alerting, logging, and APM. • Experience working in regulated environments with cloud security, compliance, governance, and audit requirements. • Strong technical leadership, communication, stakeholder management, and problem-solving skills.