Overview
Responsibilities
As a ML System Engineer on the AI & ML Platform’s Inference team, you will design and optimize large-scale model serving systems end-to-end. You will have the chance to own everything from distributed infrastructure (global KV cache, continuous batching, load balancing, auto-scaling) to deep low-level optimizations (GPU kernels, quantization, speculative decoding).
In this role, you are expected to:
Architect and implement scalable distributed infrastructure for model serving (load balancing, auto-scaling, batch scheduling, global KV cache).
Optimize latency and throughput of model inference under real production workloads.
Build reliable, high-concurrency serving systems that serve billions of requests reliably
Benchmark, fine-tune, and accelerate inference engines.
Create robust CI/CD infrastructure for seamless model deployment and inference engine updates.
Partner with senior ML engineers to finetune and deploy open-source LLMs
Compensation
At Atlassian, we strive to design equitable, explainable, and competitive compensation programs. We follow consistent hiring practices and account for each candidate's skills, knowledge, and experience when setting base pay within the range.
Please visit go.atlassian.com/payzones for more information on which locations are included in each of our geographic pay zones. However, please confirm the zone for your specific location with your recruiter.
This role may also be eligible for benefits, bonuses, commissions, and equity.
Pay Ranges
In The United States, we have three geographic pay zones. For this role, our current base pay ranges for new hires in each zone are:
Zone A: $178,200 - $232,650
Zone B: $160,380 - $209,385
Zone C: $147,906 - $193,100
Qualifications
On your first day, we’ll expect you to have:
3+ years of software engineering experience
Deep low-level systems programming (C/C++ or Rust)
Experience with large-scale, high-concurrent production serving.
Experience with GPU inference engines (vLLM, SGLang, Triton, TensorRT-LLM, etc.).
It would be great, but not required if you have:
1+ years of system performance optimization experience
Low-level inference optimizations: GPU kernels
Algorithmic inference optimizations: quantization, speculative decoding, distillation
Experience with testing, benchmarking, and reliability of inference services.
Experience designing and implementing CI/CD infrastructure for inference.
Strong background in system optimizations: batching, caching, load balancing, parallelism.
Benefits & Perks
Atlassian offers a wide range of perks and benefits designed to support you, your family and to help you engage with your local community. Our offerings include health and wellbeing resources, paid volunteer days, and so much more. To learn more, visit go.atlassian.com/perksandbenefits.