Requisition Id 17094
Overview:The National Center for Computational Sciences (NCCS) at Oak Ridge National Lab (ORNL), which hosts several of the world's most powerful computer systems, is seeking highly qualified individuals to play a key role in improving the security, performance, and reliability of the NCCS computing infrastructure which supports multiple highly ranked Top500 Supercomputers, including the world's first exaflop system, Frontier.
The Team:The HPC Platforms design, manage, and operate Kubernetes clusters that support our supercomputing center's mission. Our primary platform is the OLCF Slate Service, a container orchestration platform built on Rancher Kubernetes Engine 2 (RKE2). Slate hosts critical operational services and user-managed persistent applications that are deployed alongside and integrated with OLCF supercomputing systems and other OLCF-managed HPC clusters.
The Role:As a Kubernetes Platform Principal Engineer / Architect and Technical Lead, you will architect, implement, and maintain the infrastructure underpinning our on-premises Kubernetes clusters, with a strong focus on scalability, reliability, and maintainability. You will lead the technical direction of our platform engineering initiatives, evaluate and integrate key technologies, and mentor a team of engineers to deliver a robust internal platform that powers development across the organization.
Major Duties/Responsibilities:- Platform Architecture & Implementation
- Lead the architecture, design, implementation, and lifecycle management of on-premises Kubernetes (RKE2) clusters.
- Evaluate, select, integrate, and standardize core platform components for networking, CI/CD, operating-system management, service mesh, and Kubernetes operators. Observability services are owned by a dedicated SRE sub-team, this role will ensure platform components integrate effectively with those services.
- Build and maintain test environments to evaluate candidate technologies for performance, functionality, operational reliability, and maintainability, particularly where integration with on-premises hardware and operating-system requirements is critical.
- Lead cluster upgrade strategy and execution, security hardening, scalability planning, and integration with monitoring and observability services.
- Infrastructure as Code (IaC) & Tooling
- Develop and maintain infrastructure, configuration, and deployment automation using technologies such as Argo CD and GitOps workflows, Puppet, Go, Python, Bash, and GitLab CI.
- Establish and promote reusable automation patterns, configuration standards, and engineering practices for operating Kubernetes clusters at scale.
- Support adoption and effective use of internally developed Kubernetes operators and serve as a secondary maintainer for designated controllers.
- Internal Developer Platform & Enablement
- Collaborate on the design and delivery of a next-generation internal developer platform, informed by capabilities offered by tools such as Backstage and AWS Proton, to improve developer productivity, platform consistency, and security.
- Partner with the cybersecurity team to define secure container-image baselines and automate image and golden base image workflows.
- Engage with development teams to understand platform needs and tailor the cluster experience to meet evolving requirements.
- Technical Leadership & Mentorship
- Provide architectural guidance, code reviews, and pair programming support to a team of 8-12 engineers.
- Mentor engineers in Kubernetes, platform engineering, automation, and operational best practices.
- Contribute to onboarding, team documentation, and process improvement initiatives.
- Serve as a primary technical resource for Kubernetes platform architecture, operations, and engineering practices across the organization.
- Collaboration
- Partner with cybersecurity, SRE, and development teams to ensure the platform meets security, compliance, reliability, and usability requirements.
- Lead or contribute to cross-functional initiatives involving platform enhancements, cluster lifecycle automation, and service integration.
- Represent the HPC Platforms team in engagements with vendors and internal and external collaborators and partners.
What Sets This Role Apart- Provide hands-on technical leadership for the full lifecycle of on-premises Kubernetes infrastructure, from operating-system and hardware integration through cluster services and deployment automation.
- Shape a platform centric engineering approach that balances security, developer experience, reliability, and operational scalability.
- Strong mentorship and team enablement focus; guiding engineers while staying hands-on with operations, architecture, and implementation.
Basic Qualifications:- BS degree in computer science or related field and a minimum of 8 years of platforms engineering experience. At least 5 years of Kubernetes experience. An equivalent combination of education and experience may be considered.
- Demonstrated experience designing, implementing and operating highly available Kubernetes platforms or services in mission critical production environments.
- Strong knowledge of Linux(RHEL)/UNIX system fundamentals, including experience managing Linux-based operating systems in heterogeneous environments.
- Strong understanding of networking principles and common network protocols, including Linux and Kubernetes networking.
- Experience developing and maintaining automation, infrastructure, configuration-management, or deployment tooling using Bash and one or more high-level languages, such as Python or Go.
- Experience with infrastructure-as-code, automated configuration management, and CI/CD or GitOps practices, using tools such as Puppet, Ansible, GitLab CI, Argo CD, or comparable technologies.
- Ability to identify technical requirements; define, plan, and implement effective solutions; and independently prioritize and complete complex projects.
- Strong communication, collaboration, and technical leadership skills, including the ability to provide design guidance, code reviews, and mentorship to other engineers.
Preferred Qualifications:- Experience operating Kubernetes platforms in on-premises, bare metal, or HPC environments.
- Experience with Rancher Kubernetes Engine 2 (RKE2), Helm, Kubernetes operators/controllers, service-mesh technologies, and GitOps-based deployment workflows.
- Experience implementing systems-level security controls, such as SELinux, and applying container, operating-system, and Kubernetes security best practices.
- Experience with secure container image pipelines, image hardening, vulnerability remediation, and base/golden image lifecycle management.
- Experience contributing to open-source communities, including contributions or patches accepted upstream.
- Experience supporting internal developer platforms or developer self-service capabilities.
- Experience supporting large-scale scientific computing, high-performance computing, or research computing environments.
Special Requirements:- Visa sponsorship: Visa sponsorship is not available for this position.
- Security, Credentialing, and Eligibility Requirements: Q Clearance: This position requires the ability to obtain and maintain a clearance from the Department of Energy. As such, this position is a Workplace Substance Abuse (WSAP) testing designated position. WSAP positions require passing a pre-placement drug test and participation in an ongoing random drug testing program.
This position will remain open for a minimum of 5 days after which it will close when a qualified candidate is identified and/or hired.
We accept Word (.doc, .docx), Adobe (unsecured .pdf), Rich Text Format (.rtf), and HTML (.htm, .html) up to 5MB in size. Resumes from third party vendors will not be accepted; these resumes will be deleted and the candidates submitted will not be considered for employment.