5+ years designing and operating large-scale Kubernetes infrastructure in production.
Bachelor's or Master's degree in relevant engineering fields.
Experience with production-level large-scale network services.
In-depth knowledge of Kubernetes internals including API server and scheduler.
Strong command of CNI, load balancing, and cloud networking.
Proficient with AWS, Terraform, Helm, and Ansible.
Programming skills in Go or Python for infrastructure automation.
Excellent debugging skills in distributed systems.
Responsibilities
Own architecture of multi-cluster Kubernetes fleet including upgrades and control plane management.
Extend Kubernetes functionality with custom controllers and CRDs.
Design GPU scheduling and capacity strategy for multi-tenant environments.
Build autoscaling solutions that respond to inference traffic.
Manage Kubernetes network data plane including DNS and load balancing.
Design cross-region and cross-cluster connectivity with a service mesh.
Define and lead the development of SLOs for critical systems.
Benefits
Flexible working hours.
Daily lunch and dinner provided; unlimited snacks and beverages.
Supportive and highly collaborative work environment.
Health check-up support and top-tier equipment/hardware support.
Opportunity to be at the forefront of AI infrastructure innovation.
Competitive compensation package including health insurance and startup equity.
Full Job Description
About the job
FriendliAI is looking for a Cloud Infrastructure Engineer to own the architecture and evolution of the cluster platform behind our GPU-accelerated AI inference cloud. As a Software Engineer, Cloud Infrastructure, you will design how our clusters are built and connected, extend Kubernetes where its defaults fall short, and own the network path that inference traffic depends on.
Inference is an unforgiving workload for Kubernetes. Traffic is bursty and latency-sensitive, GPU capacity is scarce and inelastic, tenants must stay isolated, and multi-node serving depends on the network holding up under sustained load. This is a hands-on architecture role for an engineer who has already run large clusters in production and wants to push them further.
Key Responsibilities
Cluster Architecture
Own the architecture of our multi-cluster, multi-tenant Kubernetes fleet across both managed and self-managed clusters: cluster topology, control plane and etcd lifecycle, and zero-downtime upgrades.
Extend Kubernetes with custom controllers, operators, and CRDs so platform behavior is encoded in software rather than runbooks.
Design GPU scheduling and capacity strategy, including topology-aware placement, node pools, priority and preemption, and quota across tenants.
Build autoscaling that matches inference traffic: queue-driven pod scaling, node autoscaling, scale-to-zero, and cold-start reduction.
Networking
Own the Kubernetes network data plane: CNI, IPAM, DNS, ingress, and L4/L7 load balancing.
Design cross-AZ, cross-region, and cross-cluster connectivity, and operate the service mesh for routing, mTLS, and traffic policy.
Debug production network issues (packet loss, conntrack exhaustion, MTU mismatches, DNS latency, load balancer behavior) and drive permanent fixes.
Reliability & Collaboration
Define SLOs for platform-critical systems and lead post-incident hardening.
Deliver infrastructure as code with Terraform, Helm, and GitOps.
Partner with the inference engine, platform, SRE, and security teams to turn serving requirements into platform capabilities.
Qualifications
5+ years designing, building, and operating large-scale Kubernetes infrastructure in production.
Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent.
Proven experience operating large-scale, high-traffic network services in production.
Deep understanding of Kubernetes internals: API server, scheduler, controller loops, kubelet, and etcd.
Strong command of Kubernetes and cloud networking: CNI, kube-proxy/eBPF datapaths, DNS, load balancing, service mesh, and VPC routing.
Proficiency with AWS, Terraform, Helm, and Ansible.
Programming skills in Go or Python, with the ability to build infrastructure tooling and automation.
Strong debugging skills across distributed systems, containers, and the Linux networking stack.
Clear written and verbal communication, including the ability to document architectural decisions for other engineers.
Preferred Experience
Large-scale Kubernetes operations in a high-traffic domain such as gaming, e-commerce, or public cloud.
Cilium and eBPF, including kube-proxy replacement or upstream contributions.
Cluster provisioning and lifecycle management with Kubespray or similar Ansible-based tooling.