About the roleAs a Member of Technical Staff, you will build the systems that schedule, route, and coordinate AI workloads across Gimlet's infrastructure.
Different stages of an inference pipeline may run on different hardware, scale independently, and exchange state across the system. Your work will determine how those workloads are placed, coordinated, routed, recovered, and operated in production.
You will work across scheduling, orchestration, control planes, APIs, and fault tolerance. You will design systems that make distributed infrastructure easier to operate, enable workloads to run reliably across a heterogeneous fleet, and partner with compiler, ML systems, networking, and infrastructure engineers to connect the full execution stack.
What success looks likeIn the first 12-18 months, you will:
- Build scheduling and orchestration systems for heterogeneous compute
- Design systems that manage independently scalable stages of distributed inference pipelines
- Improve the reliability and fault tolerance of production AI infrastructure
- Develop control planes and APIs that simplify how workloads are deployed and managed
- Improve resource management and scheduling as Gimlet expands across new accelerator types, nodes, and data centers
- Help the platform scale across additional hardware, nodes, and data centers
You may be a good fit if you have- Experience building or operating distributed systems in production
- Strong software-engineering and systems fundamentals
- The ability to reason about concurrency, consistency, failure modes, and system tradeoffs
- Experience with scheduling, resource management, RPC, or asynchronous messaging
- A bachelor's degree in a relevant field or equivalent practical experience
Strong candidates may also have- Experience with Kubernetes or Kubernetes-adjacent systems beyond basic usage
- Experience designing service-oriented architectures using RPC or asynchronous messaging
- Familiarity with scheduling, queues, or resource management systems
- Experience building reliable APIs and operating systems under high load
- Software development experience in languages commonly used for systems development (e.g., Go, C++, Python)