About the TeamThe AI Model Serving team is the engine behind every production Workday agent and machine learning use case. We own the services that power all production AI workloads, acting as both the gateway to vendor-hosted LLMs (GCP, AWS Bedrock, Gemini) and the primary platform where Workday hosts and scales its internal models.
We operate at scale, hosting thousands of traditional ML models across sharded Ray Serve clusters and maintaining Workday's production model registry. Our platform consistently handles ~2,000 requests per second, peaking at over 10,000 RPS in our largest clusters.
In the year ahead, our engineering roadmap is highly ambitious. We are focused on:
Scaling Architecture: Upgrading our systems to seamlessly support 20+ new AI agents going into production.
Hosting Open-Weight LLMs: Designing the infrastructure to host and tune open-source LLMs directly within our stack.
Performance & Reliability: Architecting optimizations to drive down core latency while maintaining the rock-solid stability our high-throughput production systems demand.
Enterprise Governance: Hardening our unified vendor interface and implementing advanced cost-governance controls.
Our culture is built on focus, camaraderie, and high performance. We are a friendly, dedicated group that takes pride in building and operating one of the most heavily used services at Workday. If you are energized by working on the infrastructure that sits at the very heart of Workday's AI strategy, this is the team for you.
About the RoleAs either a Senior Software Development Engineer on the AI Model Serving team, you will be a technical leader who helps shape the vision and direction of the platform alongside the engineering manager. You will play a central role in making critical design decisions, driving outcomes across the team, and setting a positive and inclusive team culture.
Your work will directly impact Workday's ability to serve AI at scale - from traditional ML models to the latest large language models powering Workday's agents.
Key Responsibilities:- Lead the team technically by making critical design decisions that drive performance, reliability, and scalability across the platform.
- Design, implement, and maintain large-scale systems that enable moving ML models to production.
- Write design documents to build consensus for new system components and enhancements to existing components.
- Evaluate and uptake new technologies made available within Workday and across the broader industry.
- Troubleshoot, improve, and scale continuous integration software pipelines.
- Develop relationships with software engineers, machine learning engineers, and data scientists on partner teams.
- Respond to alerts and debug production issues to maintain platform health and reliability.
- Review pull requests and enforce consistency, performance, readability, and security across code bases.
- Develop documentation to share knowledge with other engineers.
About YouAbout YouBasic Qualifications- 6+ years of related work experience in software development, with a focus on building and operating large-scale distributed systems.
- Bachelor's degree in Computer Science, Engineering, or a related technical field (or equivalent practical experience).
Other Qualifications- Software Development and Distributed Systems: Deep experience designing, building, and scaling production-grade distributed systems. You understand the full software development lifecycle - from coding standards and testing to code reviews, source control, and deployment, and can apply that knowledge to complex, high-throughput platforms.
- Python: Deep proficiency in Python, with extensive experience writing production-level code and building systems in Python-based frameworks.
- Kubernetes & GPU Infrastructure: Deep hands-on experience deploying and scaling workloads on Kubernetes, with a specific focus on GPU resource management. You understand how to optimize GPU utilization for hosting and tuning smaller open-weight LLMs using modern inference engines (e.g., vLLM, TGI). Familiarity with GPU memory constraints, serving tuned models (e.g., LoRA), and autoscaling hardware metrics.
- LLMs and Traditional ML Models: Familiarity with both large language models and traditional ML models, including how they are served, scaled, and monitored in production. You understand the operational differences and can design abstractions that serve both effectively.
- Observability: You can design and maintain monitoring strategies that provide clear insight into system health, performance, and cost.
- Communication: Excellent written and verbal communication skills, including the ability to write clear design documents, articulate complex technical ideas, and build consensus across teams.
- Mentorship: A collaborative approach to engineering, with experience mentoring other engineers and fostering an inclusive team environment.
Workday Pay Transparency StatementThe annualized base salary ranges for the primary location and any additional locations are listed below. Workday pay ranges vary based on work location. As a part of the total compensation package, this role may be eligible for the Workday Bonus Plan or a role-specific commission/bonus, as well as annual refresh stock grants. Recruiters can share more detail during the hiring process. Each candidate's compensation offer will be based on multiple factors including, but not limited to, geography, experience, skills, job duties, and business need, among other things. For more information regarding Workday's comprehensive benefits, please click here.
Primary Location: USA.CO.Boulder
Primary Location Base Pay Range: $156,000 USD - $234,000 USD
Additional US Location(s) Base Pay Range: $148,200 USD - $264,000 USD
Additional Considerations:
The application deadline for this role is the same as the posting end date stated as below:
10/05/2026
Our Approach to Flexible WorkWith Flex Work, we're combining the best of both worlds: in-person time and remote. Our approach enables our teams to deepen connections, maintain a strong community, and do their best work. We know that flexibility can take shape in many ways, so rather than a number of required days in-office each week, we simply
spend at least half (50%) of our time each quarter in the office or in the field with our customers, prospects, and partners (depending on role). This means you'll have the freedom to create a flexible schedule that caters to your business, team, and personal needs, while being intentional to make the most of time spent together. Those in our remote "home office" roles also have the opportunity to come together in our offices for important moments that matter.