About the TeamData Platform and Observability team is based in Pleasanton and Atlanta in the US, Dublin in Ireland and Chennai in India. Our focus is on the development of large scale distributed data systems to support critical Workday products and provide real-time insights across Workday's platforms, infrastructure and applications.
The team provides platforms that process 100s of terabytes of data that enable core Workday products and use cases like core HCM, Fins, AI/ML skus, internal data products and Observability. If you enjoy writing efficient software or tuning and scaling large distributed systems you will enjoy working with us.
Do you want to tackle exciting challenges at massive scale across private and public clouds for our 10000+ global customers? Do you want to work with world class engineers and facilitate the development of the next generation Distributed systems platforms? If so, we should chat.
About the RoleThe Messaging, Streaming and Caching team is a full-service Distributed Systems Engineering team. We architect and provide async messaging, streaming, and NoSQL platforms and solutions that power the Workday products and SKUs ranging from core HCM, Fins, Integrations, and AI/ML. We develop client libraries and SDK's that make it easy for teams to build Workday products. We develop automation to deploy and run hundreds of clusters, and we also operate and tune our clusters as well. As a team member you will play a key role in improving our services and encouraging their adoption within Workday's infrastructure both in our private cloud and public cloud. As a member of this team you will design and build new capabilities from inception to deployment to exploit the full power of the core middleware infrastructure and services, and work hand in hand with our application and service teams!
Primary Responsibilities
- Design, build, and enhance critical distributed services, including Kafka, Redis, RabbitMQ etc.
- Design, develop, build, deploy and maintain core distributed services using a combination of open source and proprietary stacks across diverse infrastructure environments (Kubernetes, OpenStack, Bare Metal, etc.)
- Design and develop core software modules for streaming, messaging and caching.
- Build observability modules, alerts and automation for Dashboard lifecycle management for the distributed services.
- Build, deploy and operate infrastructure components in production environments.
- Champion all aspects of streaming, messaging and caching with a focus on resiliency and operational excellence.
- Evaluate and implement new open-source and cloud-native tools and technologies as needed.
- Participate in the on-call rotation to support the distributed systems platforms.
- Manage and optimize Workday distributed services in AWS, GCP & Private cloud env.
About YouBasic Qualification
- 12+ years experience in software development engineering.
- 6+ years experience specifically focused on designing, building, and operating distributed systems like Redis, Kafka, RabbitMQ or NoSQL solutions.
- 5+ experience in designing and implementing complex distributed system architectures, evidenced by successful deployment of systems with high availability (e.g., 99.99% uptime) and fault tolerance.
- 8+ years experience with at least two of the following programming languages (e.g., Java, Python, Go, C/C++), including experience in writing production-level code for distributed systems.
- Expertise with configuration management using Chef and service deployment on Kubernetes via Helm and ArgoCD
- Bachelor's degree in a relevant field such as Computer Science, Engineering, or a related discipline; a Master's degree (e.g., MS in Computer Science, Distributed Systems, or related field) is strongly preferred or equivalent practical experience.
Other Qualification
- Expert-level ability in Algorithmic Thinking, including CAP theorem, queuing theory, consensus protocols etc, to architect highly efficient and scalable solutions for complex distributed systems implementations.
- Deep expertise in API and Client Library Development, including understanding of RESP protocol, Kafka wire protocol etc and extensive experience in designing and building API layer as well as client libraries.
- Experience building cloud native controllers for distributed systems. Familiarity with operator like Strimzi would be a bonus
- Strong understanding of modern Code Testing methodologies like consistency / linearizability testing, and experience in leading chaos and and fault injection testing strategies.
- Deep understanding of Distributed Systems Software principles, including fault tolerance, high availability, and extensive experience in replication / sharding techniques
- Proven ability to design and implement High Availability strategies for critical distributed systems, including global replication with 99.99% SLOs.
- Extensive experience with Large Scale Data Processing technologies and frameworks such as Kafka/Redis/RabbitMQ/Spark/Flink etc within complex distributed architectures.
- Strong understanding of Object-Oriented Design (OOD) principles and architectural patterns for building highly scalable and maintainable distributed systems.
- Extensive experience with SCM and CI/CD tools such as Git, Jenkins, Harness etc and establishing best practices for collaborative distributed development workflows.
- Strong understanding of System Security principles and best practices relevant to securing complex distributed environments, including implementing authentication and authorization modules.
- Experience learning complex open source service internals via code inspection.
- Proven ability to lead Team Collaboration within and across distributed software development teams and drive architectural direction.
- Strong skills in creating Technical Writing Documentation and presenting to senior leadership as well as architects within the company.
Workday Pay Transparency StatementThe annualized base salary ranges for the primary location and any additional locations are listed below. Workday pay ranges vary based on work location. As a part of the total compensation package, this role may be eligible for the Workday Bonus Plan or a role-specific commission/bonus, as well as annual refresh stock grants. Recruiters can share more detail during the hiring process. Each candidate's compensation offer will be based on multiple factors including, but not limited to, geography, experience, skills, job duties, and business need, among other things. For more information regarding Workday's comprehensive benefits, please click here.
Primary Location: USA.CA.Pleasanton
Primary Location Base Pay Range: $222,900 USD - $334,300 USD
Additional US Location(s) Base Pay Range: $187,100 USD - $334,300 USD
Our Approach to Flexible WorkWith Flex Work, we're combining the best of both worlds: in-person time and remote. Our approach enables our teams to deepen connections, maintain a strong community, and do their best work. We know that flexibility can take shape in many ways, so rather than a number of required days in-office each week, we simply
spend at least half (50%) of our time each quarter in the office or in the field with our customers, prospects, and partners (depending on role). This means you'll have the freedom to create a flexible schedule that caters to your business, team, and personal needs, while being intentional to make the most of time spent together. Those in our remote "home office" roles also have the opportunity to come together in our offices for important moments that matter.