Mission:This role is for a seasoned engineer with practical experience making LLMs fast, efficient, and reliable in production. This individual will work on an inference serving software stack from a trained model to a high-throughput delivery of output tokens. This will require experience on understanding/working with model internals, runtime software, and the supporting GPU & accelerator hardware underpinning it all.
Location: This role will be based in one of our three hiring hubs: the Dallas, San Francisco, or New York City area. The person hired for this role must be based in one of these three areas. You'll have the flexibility to work remotely while we establish our local Groq office, with the expectation that this role will transition to onsite once the office opens.
Responsibilities & opportunities in this role:- Build and optimize serving systems that handle high request volumes with low, predictable latency
- Apply model compilation and graph optimization techniques, and benchmark carefully to use them only where they deliver real gains
- Implement and tune batching, caching, scheduling, and memory management strategies for large models
- Apply reduced-precision and quantization methods
- Distribute models across multiple accelerators and nodes, balancing throughput & latency
- Profile end-to-end performance, identify bottlenecks, and fix them-ranging from low-level kernels to request routing-leveraging both lab/development environments and live telemetry from production systems.
- Partner with infrastructure and FDE teams to bring new models into production quickly
- Partner with vendors along the inference serving path in support of optimizing the stack
Ideal candidates have/are:- BS / MS / PhD in CS, CE, EE, or equivalent depth from industry
- 5+ years shipping performance-critical products, with some portion of that serving models at scale
- A strong applied understanding of transformer architectures
- Fluency in at least one systems language and one high-level language
- A rigorous, measurement-driven approach to performance work
- Clear communication about tradeoffs to both technical and business stakeholders
Ways to stand out:- Experience writing or tuning custom kernels.
- Familiarity with non-GPU or specialized inference hardware.
- Contributions to open-source serving or compiler projects.
CompensationGroq is committed to providing competitive compensation through our Total Cash philosophy, which incorporates potential bonus value directly into base pay. The total cash salary range for this position, which is inclusive of the potential bonus value, is $341,400 - $401,600, with individual placement determined by your geographic location, experience, skills, and alignment with internal compensation standards. This range is specific to candidates located in the United States. Compensation for international candidates will vary based on local market dynamics. Beyond cash compensation, Groq also offers a Long-Term Incentive (LTI) Program and a robust suite of employee benefits.
#LI-MS1