Full Job Description
We are looking for an experienced Full-Stack Software Engineer to design and implement the Infrastructure Services that power our GPU-as-a-Service platform. You will build the control plane that turns high-level API calls into real infrastructure actions - enrolling bare-metal servers, provisioning them, and assembling them into multi-tenant Kubernetes clusters on high-performance hardware.
You will own the full lifecycle of the infrastructure-level services-from the Server and MachineType APIs down to the provisioning workflows and reconciliation loops that keep the platform's view of hardware consistent with physical reality.
Key Responsibilities
• Infrastructure API Design: Design, build, and maintain the versioned REST and gRPC APIs for bare-metal server lifecycle, MachineType definitions, and cluster CRUD operations.
• Provisioning & Workflow Development: Develop the asynchronous workflows that drive server enrollment, inspection, OS provisioning, and cluster bring-up, exposing durable status to callers.
• Full-Stack Implementation: Implement and maintain the consoles and interfaces that visualize hardware inventory, provisioning progress, and cluster health.
• System Reliability: Design the error-handling models, idempotency guarantees, and reconciliation loops necessary to manage long-running provisioning operations reliably.
Qualifications
• API Development: Strong experience designing RESTful APIs or gRPC services. You understand API versioning and gateway patterns.
• Language Stack: Proficiency in Go (preferred for backend/Kubernetes ecosystem) and modern TypeScript/React (for the Console).
• Kubernetes Knowledge: Deep understanding of Kubernetes primitives and controller/reconciler patterns. You will be interacting with systems like k0rdent, Metal3, and Cluster API to translate high-level API calls into infrastructure actions.
• Bare-Metal Provisioning: Hands-on experience with bare-metal provisioning flows - BMC/Redfish, PXE/iPXE, image management, and hardware inspection.
• Asynchronous Systems: Experience building workflow-driven or event-driven systems (e.g., Temporal) where operations are long-running and state must remain consistent across retries and failures.
Preferred Qualifications
• State Reconciliation: Experience building informers or reconciliation bridges that keep an external datastore consistent with Kubernetes resource state.
• Multi-Tenancy: Experience building platforms where strict data and network isolation between tenants is required.
• Infrastructure-as-Code: Familiarity with Terraform/OpenTofu and GitOps-driven configuration (ArgoCD or Flux).
• Hardware Domain: Familiarity with GPU server hardware, DPUs/NICs, and high-performance datacenter fabrics.
Additional Information
What does Mirantis offer you?
- Work with an established Silicon Valley leader in the cloud infrastructure industry;
- Work with exceptionally passionate, talented and engaging colleagues, helping Fortune 500 and Global 2000 customers implement next-generation cloud technologies;
- Be a part of cutting-edge, open-source innovation;
- Thrive in the high-energy environment of a young company where openness, collaboration, risk-taking, and continuous growth are valued;
- Professional development and training;
- Attend conferences and working groups;
- Company outings, happy hours, hackathons, and tech talks;
- Receive a competitive compensation package with a strong benefits plan.