Position OverviewWe are seeking a Principal Kubernetes Control Plane Engineer to architect the foundational control plane for our AI-native NeoCloud platform. You will be responsible for building highly available, automated systems to manage the lifecycle of thousands of Kubernetes control planes, ensuring zero-downtime upgrades and seamless multi-cluster federation. This role is central to our mission of providing a "NeoCloud" experience, requiring a strategic mindset to balance complex distributed systems design with robust, production-grade engineering. You will lead the development of the "chassis" that supports our high-performance AI workloads, ensuring our infrastructure is scalable, resilient, and ready for enterprise-grade demands.
Key Responsibilities- Design and implement scalable control plane architectures using tools like Cluster API to enable fleet-wide automation.
- Develop custom Kubernetes Operators and Custom Resource Definitions (CRDs) to streamline lifecycle management and orchestration.
- Optimize etcd performance and ensure rock-solid state management across distributed cloud regions.
- Implement secure, multi-tenant isolation mechanisms at the API server layer to support enterprise-grade security and tenancy.
- Architect solutions for zero-downtime upgrades and multi-cluster federation to provide a seamless user experience for our AI customers.
- Collaborate with AI Scheduling and Fabric engineering teams to ensure the control plane integrates tightly with GPU orchestration layers.
- Lead design reviews and mentor engineers to uphold high standards of distributed systems architecture across the team.
- Drive platform reliability by identifying and mitigating failure modes in the control plane chassis before they impact customer training or inference jobs.
Qualifications- Bachelor's or Master's degree in Computer Science, Electrical Engineering, or related field.
- 8+ years of software engineering experience, with deep, hands-on expertise in Go.
- Extensive experience contributing to or extending the Kubernetes ecosystem (e.g., Operators, API Server, etcd, Cluster API).
- Proven track record of operating, debugging, and scaling large-scale distributed systems in production environments.
- Deep understanding of multi-tenancy models, cluster federation, and API governance in cloud-native environments.
- Familiarity with infrastructure automation (e.g., Terraform, CI/CD pipelines) and managed cloud services.
- Strong technical leadership skills; ability to influence architectural decisions and align cross-functional teams.
- Excellent communication skills, with the ability to translate complex system requirements into manageable engineering milestones.
- Experience working in high-velocity, high-growth engineering environments is highly preferred.