This position requires access to customer data; Must be a U.S. citizen. SAP NS2 does not offer Visa sponsorships for this role. All internals must have manager's approval to transfer. Role OverviewSAP NS2 is seeking an accomplished Director, DevOps & Cloud Platform Engineering to lead the strategy, architecture, delivery, and operational excellence of our secure cloud infrastructure and automation capabilities. This leader will own a critical function responsible for enabling highly available, scalable, compliant, and automation-driven cloud environments that support SAP products and services across a diverse customer base, including local, state, and federal government agencies, as well as private sector organizations supporting government-related missions.
This role is designed for a seasoned technology leader with deep expertise in multi-cloud infrastructure, platform engineering, DevSecOps, site reliability practices, and secure operations at scale. The successful candidate will lead a high-performing engineering organization responsible for designing, building, and operating provider-agnostic infrastructure solutions across AWS, Azure, and GCP, while driving standardization, resilience, and continuous improvement across the platform lifecycle.
The ideal candidate combines executive-level leadership, strong engineering judgment, operational discipline, and hands-on technical credibility. This individual will play a key role in advancing SAP NS2's cloud maturity by scaling infrastructure automation, strengthening reliability and security practices, and ensuring mission-critical environments meet demanding performance, compliance, and customer expectations.
Key Responsibilities- Lead, mentor, and scale a high-performing DevOps and cloud platform engineering organization, fostering a culture of technical excellence, accountability, innovation, and continuous improvement.
- Define and execute the strategic roadmap for cloud infrastructure, platform automation, operational maturity, and engineering standardization across a multi-cloud ecosystem.
- Serve as the senior technical and people leader for infrastructure engineering, balancing long-term platform strategy with near-term operational execution.
- Oversee the architecture, deployment, and lifecycle management of secure, scalable, resilient infrastructure using Infrastructure as Code, with a strong emphasis on Terraform, reusable modules, policy-driven automation, and platform consistency.
- Drive the design and operational governance of provider-agnostic infrastructure patterns across AWS, Azure, and GCP to support secure, repeatable, and compliant deployments.
- Lead the evolution of configuration management and post-provisioning automation practices using Ansible and related tooling to support customer environments at scale.
- Establish and enforce engineering best practices for GitLab-based source control, branching strategy, peer review, release governance, CI/CD, and change management across multiple repositories and teams.
- Provide leadership for the deployment, operation, and continuous improvement of containerized and orchestrated platforms, including Docker and Kubernetes, with a focus on reliability, scalability, maintainability, and security.
- Partner closely with product engineering, security, compliance, architecture, operations, and customer-facing teams to deliver infrastructure solutions aligned with business priorities, customer commitments, and regulatory requirements.
- Champion DevSecOps and security-by-design principles, ensuring security, governance, and compliance controls are embedded into infrastructure architecture, automation pipelines, and operational processes.
- Drive operational excellence through mature incident response, root cause analysis, problem management, service restoration, and post-incident improvement practices.
- Provide executive oversight for troubleshooting and resolving complex infrastructure, platform, networking, and connectivity issues across hosted and customer-connected environments.
- Ensure robust observability, monitoring, logging, alerting, and telemetry capabilities are implemented and continuously improved across a multi-tenant, multi-cloud environment.
- Lead efforts to improve service reliability, platform resilience, disaster recovery readiness, and operational performance for mission-critical workloads.
- Oversee the development and maintenance of high-quality technical documentation, architecture artifacts, standard operating procedures, governance materials, knowledge base content, and technical advisories.
- Identify and responsibly implement opportunities to leverage generative AI and intelligent automation to improve engineering productivity, consistency, speed, and operational effectiveness within approved security and compliance boundaries.
- Manage organizational priorities, staffing, delivery planning, cross-functional dependencies, and execution risk to ensure successful delivery of both strategic initiatives and day-to-day operational commitments.
Knowledge and Skills- Proven leadership experience managing DevOps, cloud infrastructure, platform engineering, or SRE teams in complex enterprise or regulated environments.
- Deep expertise in AWS foundational and advanced services, including EC2, S3, IAM, Route 53, VPC, and related networking, security, and automation capabilities.
- Deep expertise in Azure foundational services, including Virtual Networks, Application Gateway, Storage Accounts, Virtual Machines, Load Balancer, and Resource Groups.
- Strong expertise in GCP foundational services, including Projects, Compute Engine, GKE, Cloud Storage, and VPC.
- Expert-level proficiency in Terraform, including module architecture, reusable design patterns, state management, policy integration, and multi-cloud deployment strategies.
- Strong experience leading CI/CD transformation and infrastructure automation maturity across engineering organizations.
- Strong scripting and automation skills using Python, Bash, and related tooling.
- Experience with RBAC, identity provisioning, secrets management, and access governance, including integration with enterprise identity and security controls.
- Strong understanding of platform engineering, SRE principles, and service reliability practices, including SLAs, SLOs, error budgets, and operational metrics.
- Strong investigative and troubleshooting capabilities across cloud infrastructure, distributed systems, Linux platforms, and network configurations.
- Experience with observability, centralized logging, telemetry, and log processing in cloud-native and hybrid environments.
- Strong knowledge of networking fundamentals, including IP routing, subnetting, DNS, load balancing, firewalls, and connectivity troubleshooting.
- Strong Linux systems expertise, including deployment, hardening, performance tuning, patching, and troubleshooting.
- Familiarity with enterprise workflow and service management platforms such as ServiceNow and Jira.
- Excellent written and verbal communication skills, with the ability to influence and communicate effectively across executive, technical, operational, and customer stakeholder groups.
Minimum Qualifications- Bachelor's degree in Computer Science, Engineering, Information Systems, or equivalent practical experience.
- 10+ years of experience in cloud infrastructure, DevOps, site reliability engineering, platform engineering, or related technical disciplines.
- 5+ years of experience leading or managing technical teams in a DevOps, infrastructure, cloud engineering, or platform operations environment.
- Demonstrated hands-on experience with technologies such as Terraform, Ansible, Python, CI/CD platforms, Vault, Kubernetes, Docker, and identity/access management solutions.
- Strong experience administering and securing Unix/Linux systems in enterprise or cloud environments.
- Strong experience with networking, distributed systems, and cloud operations in highly available production environments.
- Demonstrated success leading teams responsible for secure, scalable, and highly automated infrastructure platforms.
Preferred Qualifications- Experience leading teams responsible for large-scale, distributed, multi-cloud platforms supporting mission-critical workloads.
- Demonstrated success building and scaling infrastructure automation and platform engineering programs in regulated, compliance-driven, or security-sensitive environments.
- Experience supporting government, national security, defense, or critical infrastructure customers.
- Strong background in incident command, operational excellence, service reliability, and continuous service improvement.
- Experience implementing DevSecOps controls, compliance automation, and policy-as-code in cloud environments.
- Ability to operate effectively as both a strategic leader and a technically credible engineering partner.
- Proven track record of driving standardization, automation, resilience, and engineering maturity across infrastructure operations.
- Strong problem-solving ability, sound judgment, and a high degree of ownership, accountability, and execution focus.
Requisition ID: 459620 | Work Area: Software-Design and Development | Expected Travel: 0 - 10% | Career Status: Management | Employment Type: Regular Full Time | Additional Locations: #LI-Hybrid
Requisition ID: 459620
Posted Date: Sep 2, 2026
Work Area: Software-Design and Development
Career Status: Management
Employment Type: Regular Full Time
Expected Travel: 0 - 10%
Location: