Workday

Principal Distributed Systems Engineer - Observability

Workday$222K — $334K *
Enterprise Technology
11 - 15 years of experience
Job Overview by Ladders

Qualifications

  • 14+ years of software development experience.
  • 6+ years focused on complex distributed systems design and operation.
  • 8+ years proficiency in at least two programming languages (Java, Python, Go).
  • Bachelor's degree in Computer Science or related field; Master’s preferred.
  • Expertise in algorithmic thinking and distributed systems software principles.

Responsibilities

  • Architect and build the distributed tracing platform using ClickHouse/Tempo.
  • Manage the big-data pipeline for tracing data using Kafka and Iceberg.
  • Enhance performance for data ingestion and query paths.
  • Lead the design for high availability and disaster recovery of tracing services.
  • Design multi-tenant security architecture for data access.
  • Ensure operational excellence through monitoring and alerting.
  • Evaluate and implement new technologies to improve platform capabilities.

Benefits

  • Flexible work arrangements fostering a balance of remote and in-office time.
  • Opportunities for continuous learning and professional development.
  • Access to a comprehensive benefits package, including health and wellness programs.
  • Participation in a bonus plan and annual stock grants.
  • Engagement in team-building activities and community events.
Full Job Description

About the Team

The Data Platform and Observability Engineering (DPOE) team is building Workday's next-generation, multi-petabyte scale Observability Platform. We own the libraries, distributed services, and infrastructure that power ingestion, storage, and query across the observability stack — Iceberg, ClickHouse, Tempo, Grafana, S3, Kafka, and Elasticsearch — serving traces, metrics, and logs for every workload at Workday. Our roadmap directly shapes how the company detects, diagnoses, and eventually predicts operational issues at scale.

About the Role

To own the technical vision and architecture for distributed tracing as a first-class pillar of Workday's Observability Platform, built on ClickHouse and/or Grafana Tempo, backed by a big-data pipeline (Kafka, Spark/Flink, Iceberg, Clickhouse, Tempo,S3) running on AWS. This is a hands-on, high-autonomy role for an engineer who can design and build multi-petabyte, low-latency tracing infrastructure end-to-end — and who is equally excited to help define where Observability AI goes next: using traces, logs, and metrics as the substrate for automated root-cause analysis, anomaly detection, and AI-driven incident triage.

You'll set technical direction across multiple teams, mentor senior and staff engineers, and act as the primary architect and escalation point for the tracing subsystem — from ingestion and storage design through query performance and platform reliability.

  • Architect and build Workday's distributed tracing platform on ClickHouse/Tempo, designed for multi-petabyte scale ingestion and sub-second interactive query performance.
  • Own the big-data pipeline feeding tracing data — Kafka-based ingestion, Spark/Flink stream and batch processing, and Iceberg-on-S3 storage — including schema design, partitioning, compaction, and lifecycle management.
  • Drive performance and scaling across ingestion and query paths: storage format optimization (Parquet/Iceberg), compression strategy, partitioning/indexing, and query engine tuning under real production load.
  • Lead HA/DR design for tracing services — multi-region/multi-AZ resilience, failover, backup/restore, and recovery time/point objectives appropriate to a tier-1 platform.
  • Design security architecture for the platform, including authentication/authorization (authn/authz) for multi-tenant data access across ingestion and query layers.
  • Own operational excellence for distributed tracing: monitoring, logging, alerting, capacity planning, and participation in an on-call rotation for the platform.
  • Evaluate and introduce new technologies — open source and cloud-native — that materially improve the platform's scalability, cost efficiency, or capability.
  • Shape the future of Observability AI: partner with ML/AI stakeholders to define how tracing data feeds automated anomaly detection, root-cause analysis, and AI-assisted incident management.
  • Evangelize the platform: publish best practices, mentor engineers across DPOE and partner teams, and act as a technical thought leader for the modern observability/data stack internally.
  • Operate with high autonomy in a fast-moving, ambiguous environment — setting technical direction with minimal oversight while aligning with broader platform strategy.

About You

Basic Qualification

14+ years experience in software development engineering.
6+ years experience specifically focused on designing, building, and operating complex distributed system architectures, evidenced by successful deployment of systems with high availability (e.g., 99.9% uptime) and fault tolerance.
8+ years experience with at least two of the following programming languages (e.g., Java, Python, Go), including experience in writing production-level code for distributed systems.
Bachelors degree in a relevant field such as Computer Science, Engineering, or a related discipline; a Master's degree (e.g., MS in Computer Science, Distributed Systems, or related field) is strongly preferred or equivalent practical experience.

Other Qualification

Expert-level ability in Algorithmic Thinking, including [insert specific advanced algorithms or data structures relevant to distributed systems], to architect highly efficient and scalable solutions for complex
Deep expertise in API Development, including understanding of advanced API protocols or architectural patterns
Deep understanding of Distributed Systems Software principles, like distributed consensus or fault tolerance mechanisms
Proven ability to design and implement High Availability strategies for critical distributed systems
Extensive experience with Large Scale Data Processing technologies and frameworks
Deep understanding of Large Scale Systems design principles like distributed data management or scalability strategies
Strong understanding of System Security principles and best practices relevant to securing complex distributed environments
Proven ability to lead Team Collaboration within and across distributed software development teams and drive architectural direction
Strong skills in creating Technical Writing Documentation and Presentation


Workday Pay Transparency Statement

The annualized base salary ranges for the primary location and any additional locations are listed below. Workday pay ranges vary based on work location. As a part of the total compensation package, this role may be eligible for the Workday Bonus Plan or a role-specific commission/bonus, as well as annual refresh stock grants. Recruiters can share more detail during the hiring process. Each candidates compensation offer will be based on multiple factors including, but not limited to, geography, experience, skills, job duties, and business need, among other things. For more information regarding Workdays comprehensive benefits, please .

Primary Location: USA.CA.PleasantonPrimary Location Base Pay Range: $222,900 USD - $334,300 USD


Additional US Location(s) Base Pay Range: $187,100 USD - $334,300 USD



Our Approach to Flexible Work

With Flex Work, were combining the best of both worlds: in-person time and remote. Our approach enables our teams to deepen connections, maintain a strong community, and do their best work. We know that flexibility can take shape in many ways, so rather than a number of required days in-office each week, we simply spend at least half (50%) of our time each quarter in the office or in the field with our customers, prospects, and partners (depending on role). This means you'll have the freedom to create a flexible schedule that caters to your business, team, and personal needs, while being intentional to make the most of time spent together. Those in our remote 4home office4 roles also have the opportunity to come together in our offices for important moments that matter.

About Workday

Workday, Inc. is a provider of enterprise cloud applications for finance and human resources. The Company delivers financial management, human capital management and analytics applications designed for various companies, educational institutions and government agencies. As part of its applications, the Company provides embedded analytics that capture the content and context of everyday business events, facilitating informed decision-making from wherever users are working. Its applications include Workday Financial Management, Workday Human Capital Management (HCM) and Other Applications. It also provides open, standards-based Web-services application programming interfaces, and pre-built packaged integrations and connectors. Workday, Inc. is headquartered in Pleasanton, California.
Learn more about Workday
Size
15,932 employees
Market Cap
$42.2 billion
Industry
Net Income
-$282.4 million
Founded
2005
5 Year Trend
+26.7%
Revenue
$4.3 billion
NASDAQ

Similar Jobs

More Jobs at Workday

More Enterprise Technology Jobs

Find similar Principal Distributed Systems Engineer - Observability jobs: