Job Summary
The Principal Platform Engineer will lead the architecture, development, and operation of an AI-native development platform that standardizes code generation, verification, security, deployment, and monitoring to enable safe and repeatable AI-assisted product development. This hands-on engineering and technical leadership role spans infrastructure as code, delivery pipelines, runtime architecture, identity and networking, security and compliance, observability, cost management, and AI-assisted development environments. The engineer will drive strategic architecture decisions, establish reusable platform capabilities, enforce security and quality standards, and mentor engineers while ensuring the platform improves developer productivity, reliability, and operational efficiency. The position follows a hybrid work model, requires travel based on client, team, and individual circumstances, and does not offer relocation assistance.
Key Responsibilities
• Drive platform architecture and runtime strategy, including standardization on a container-orchestrated runtime, tenant isolation across client engagements, and defining platform contracts inherited by product teams.
• Design composable infrastructure-as-code solutions at fleet scale, with clear input contracts, state boundaries, and versioned templates and images for core platform components.
• Automate environment lifecycle management so environments can be created through pull request merges and destroyed through pull request deletion, with automated verification that resources are removed.
• Design federated identity and network architectures with credential-free access, engagement-specific identities, just-in-time audited privileged access, and separate internal and external network boundaries.
• Provide standardized telemetry and cost visibility for platform services, including alerting, service-level objectives based on live data, fault injection, and per-engagement cost attribution, including AI-related spending.
• Build and maintain shared delivery pipelines and a pluggable gate framework for quality, security, and AI governance checks, with enforcement that blocks releases when required checks fail.
• Implement security and compliance as code, including repository scanning, policy enforcement, admission controls, signed and attested artifacts, audit trails, and software bills of materials covering AI components.
• Maintain security and compliance evidence supporting SOC 2 and ISO 27001 requirements.
• Package the AI-assisted development environment as a versioned platform deliverable and enforce sandbox boundaries for coding agents through default-deny network egress, short-lived credentials, and isolated shared runners.
• Measure platform performance through developer outcomes such as setup time, time to first commit, pipeline duration, gate false-positive rates, and adoption.
• Evaluate and improve platform controls when their governance burden exceeds their benefits, and validate standard development workflows by delivering software through the platform.
• Lead strategic architecture decisions concerning multi-tenancy, cloud posture, identity providers, and network boundaries through evidence-based decision documents that communicate technical trade-offs.
• Own and communicate security and reliability evidence used to assess platform performance and program outcomes.
• Serve as a technical authority on platform architecture and delivery engineering, advising senior stakeholders during design and strategy discussions.
• Translate complex technical decisions and platform risks into actionable recommendations for diverse audiences, including executive stakeholders.
• Mentor engineers and vendor partners, improving design quality, engineering practices, and production readiness.
• Collaborate with cross-functional product teams using Agile methodologies.
Required Qualifications
• 15+ years of experience in software engineering, platform engineering, infrastructure engineering, or a closely related technical field.
• Deep hands-on expertise in infrastructure as code at fleet scale using Terraform or OpenTofu with Terragrunt or equivalent technologies, including module architecture, remote state, and multi-environment topology.
• Demonstrated experience operating multi-tenant Kubernetes platforms in production, including GitOps, network and admission policies, workload identity, and tenant isolation.
• Proven experience building shared CI/CD platforms using GitHub Actions or equivalent technologies, including reusable workflow contracts, fleet-scale branch protection, and enforced quality and security gates.
• Strong hands-on knowledge of cloud identity and network security across Azure, AWS, GCP, or comparable cloud platforms, including workload identity federation, role-based access control (RBAC), private networking, web application firewalls (WAF), and DNS.
• Substantive experience implementing security and compliance as code using scanning tools such as Snyk, Checkov, and Trivy; policy engines such as OPA, Kyverno, and Azure Policy; and software supply chain controls such as signing, provenance, and software bills of materials (SBOM).
• Experience producing security and compliance control evidence for SOC 2 and ISO 27001.
• Production observability experience with OpenTelemetry or equivalent technologies, including collector topology, service-level objectives, error budgets, alert routing, incident management, and fault injection.
• Experience building and operating production application software as a software engineer using at least one of Go, Python, or TypeScript.
• Strong understanding of testing strategies, code reviews, trunk-based development, and release management.
• Excellent written and verbal communication skills in English, including the ability to explain complex technical concepts to diverse audiences and executive stakeholders.
• Experience working with Agile methodologies and cross-functional product teams.
Preferred Qualifications
• Experience building platforms for consulting, advisory, or professional services organizations, including client isolation and data residency requirements.
• Familiarity with virtual-cluster and control-plane technologies such as vcluster and Crossplane, as well as developer portals such as Backstage and Port.
• Experience operating platforms used by AI coding assistants such as Claude Code, GitHub Copilot, or similar tools, including sandboxing AI agent environments.
• Advanced certifications in cloud architecture, Kubernetes, or security, such as Azure Solutions Architect, CKA/CKS, or CISSP.
• FinOps experience, including cost allocation, showback, and cloud cost optimization.
• Demonstrated interest in developer education through internal documentation, workshops, or developer tooling communities.