About the team and the role:The Observability Platform team builds and operates the infrastructure that helps eBay teams monitor, fix, and improve the reliability of large-scale distributed systems. This platform supports the telemetry and reliability needs of thousands of microservices across eBay and operates at hyperscale, processing billions of time series and petabytes of log data using modern open-source technologies including Prometheus, ClickHouse, OpenTelemetry, and related tools.
As a Software Engineer on this team, you will design and build scalable distributed systems that power metrics, logs, traces, and related observability workflows across the stack-from ingestion and storage to query and visualization. You will partner closely with SREs, platform engineers, and service owners to solve complex reliability challenges, improve operational excellence, and help shape the future of observability at eBay.
This role offers the opportunity to work on critical systems, contribute to open-source technologies, and grow through direct exposure to some of eBay's most complex infrastructure challenges. This role also includes participation in the team's on-call rotation in support of production reliability.
What you will accomplish:- Design and deliver scalable, fault-tolerant observability infrastructure that improves reliability while reducing operational overhead for platform and engineering teams
- Build and optimize high-throughput services for ingesting, transforming, storing, and querying telemetry data across logs, metrics, and traces
- Strengthen the resilience of Kubernetes-based production systems through self-healing, autoscaling, and robust operational design
- Partner with SREs, platform teams, and service owners to translate observability needs into tools and platform capabilities that improve incident response and operational excellence
- Contribute to architecture reviews, production readiness discussions, and post-incident findings to drive continuous improvement across the platform
- Expand your technical breadth by working across distributed systems, cloud-native infrastructure, and optionally user-facing observability experiences.
What you will bring:- 7+ years of experience in software engineering, infrastructure engineering, or a closely related field
- Strong programming skills in Golang or another systems-level language, with experience building reliable backend or infrastructure services
- Deep understanding of distributed systems concepts such as fault tolerance, scalability, and system reliability
- Hands-on experience deploying and operating containerized services in Kubernetes or similar cloud-native environments
- Solid understanding of observability domains including metrics, logs, and traces
- Experience with tools such as Prometheus, Grafana, OpenTelemetry, ClickHouse, or similar technologies; familiarity with time-series systems, high-throughput data pipelines, open-source infrastructure, or React/JavaScript is a plus
#LI-BB1
Additional DetailsThe base pay range for this position is expected in the range below:
$172,000 - $229,600
Base pay offered may vary depending on multiple individualized factors, including location, skills, and experience. The total compensation package for this position may also include other elements, including a target bonus and restricted stock units (as applicable) in addition to a full range of medical, financial, and/or other benefits (including 401(k) eligibility and various paid time off benefits, such as PTO and parental leave). Details of participation in these benefit plans will be provided if an employee receives an offer of employment.
If hired, employees will be in an "at-will position" and the Company reserves the right to modify base salary (as well as any other discretionary payment or compensation program) at any time, including for reasons related to individual performance, Company or individual department/team performance, and market factors.
Remote roles are not eligible for U.S. visa sponsorship.