OverviewResponsibilitiesAbout This RoleAs a
Senior ML System Engineer on the AI & ML Platform's Inference team, you will design and optimize large-scale model serving systems end-to-end. You will own everything from distributed infrastructure (global KV cache, continuous batching, load balancing, auto-scaling) to deep low-level optimizations (GPU kernels, quantization, speculative decoding).
In this role, you are expected to:- Architect and implement scalable distributed infrastructure for model serving (load balancing, auto-scaling, batch scheduling, global KV cache).
- Optimize latency and throughput of model inference under real production workloads.
- Build reliable, high-concurrency serving systems that serve billions of requests reliably
- Benchmark, fine-tune, and accelerate inference engines.
- Create robust CI/CD infrastructure for seamless model deployment and inference engine updates.
- Partner with senior ML engineers to fine-tune open-source LLMs and deploy
CompensationCompensation
At Atlassian, we strive to design equitable, explainable, and competitive compensation programs. We follow consistent hiring practices and account for each candidate's skills, knowledge, and experience when setting base pay within the range. Please visit go.atlassian.com/payzones for more information on which locations are included in each of our geographic pay zones. However, please confirm the zone for your specific location with your recruiter.This role may also be eligible for benefits, bonuses, commissions, and equity.
Pay RangesIn The
United States, we have three geographic pay zones. For this role, our current base pay ranges for new hires in each zone are:
Zone A: $206,100 - $269,075
Zone B: $185,490 - $242,168
Zone C: $171,063 - $223,332
QualificationsOn your first day, we'll expect you to have
- 5+ years of software engineering experience, 2+ years of system performance optimization experience
- Deep low-level systems programming (C/C++ or Rust)
- Experience with large-scale, high-concurrent production serving.
- Experience with GPU inference engines (vLLM, SGLang, Triton, TensorRT-LLM, etc.).
- Strong background in system optimizations: batching, caching, load balancing, parallelism.
It would be great, but not required if you have
- Low-level inference optimizations: GPU kernels
- Algorithmic inference optimizations: quantization, speculative decoding, distillation
- Experience with testing, benchmarking, and reliability of inference services.
- Experience designing and implementing CI/CD infrastructure for inference.
Benefits & PerksAtlassian offers a wide range of perks and benefits designed to support you, your family and to help you engage with your local community. Our offerings include health and wellbeing resources, paid volunteer days, and so much more. To learn more, visit
go.atlassian.com/perksandbenefits