Senior OpenTelemetry Engineer
Must Have Technical/Functional Skills
This role focuses on the development of a world-class observability platform spanning SaaS products, cloud infrastructure, network fabric, and custom applications. As a Senior OTel Instrumentation Engineer, the successful candidate will own the full signal collection layer. This includes designing and deploying Open Telemetry instrumentation across all four signal domains to ensure traces, metrics, and logs are accurate, well-attributed, and reliably directed to the AIOps Lakehouse.
The position requires close collaboration with Platform, SRE, and ML Engineering teams to define instrumentation standards, evaluate new OTel components, and deliver high-fidelity telemetry that powers anomaly detection, root cause analysis, and business observability.
Roles & Responsibilities
1. SaaS Instrumentation
• End-to-End Instrumentation: Instrument SaaS products utilizing OTel SDKs (Python, Node.js, Java, Go) to ensure spans, metrics, and structured logs are emitted at the appropriate granularity.
• Semantic Conventions: Define and enforce attribute naming conventions (e.g., service.name, tenant.id, feature.flag) aligned strictly with OTel semantic standards.
• Multi-Tenant Observability: Instrument multi-tenant surfaces to guarantee robust tenant-level observability while preventing cross-tenant data leakage.
2. Cloud Infrastructure
• Collector Fleet Management: Deploy and maintain the OTel Collector fleet across AWS, GCP, and Azure, including receiver configurations, processor pipelines, and exporter routing.
• Runtime Instrumentation: Instrument serverless (Lambda, Cloud Run) and container runtimes (EKS, GKE, AKS) utilizing auto-instrumentation where feasible, and manual instrumentation when technical nuance requires it.
• Metric Normalization: Collect and normalize cloud provider metrics (CloudWatch, Cloud Monitoring, Azure Monitor) via OTel receiver plugins.
3. Network Telemetry
• Data Collection: Gather network flow data (sFlow, NetFlow/IPFIX), SNMP traps, and BGP state via OTel-native and bridged receivers.
• Context Propagation: Correlate network events with application traces using consistent trace-context propagation across network boundaries.
• Strategy Collaboration: Partner with NetOps to define the MELT (Metrics, Events, Logs, Traces) strategy for on-prem, SD-WAN, and cloud interconnect segments.
4. Application Observability
• APM Ownership: Own Application Performance Monitoring (APM) instrumentation, including distributed tracing, custom span attributes, database query capture, and error fingerprinting.
• Asynchronous Workloads: Instrument async workloads (Kafka, SQS, Celery) ensuring W3C Trace Context propagation across message boundaries.
• SLI/SLO Definition: Define and instrument Service Level Indicators (SLIs) and Service Level Objectives (SLOs)such as latency histograms, error-rate counters, and availability gaugesqueryable directly from the Lakehouse.
5. Platform & Standards
• Internal Libraries: Build and maintain internal instrumentation libraries and language-specific wrappers that encode organizational conventions.
• Governance: Author and review instrumentation RFCs while driving the adoption of OTel semantic conventions across all engineering teams.
• Data Alignment: Collaborate with the AIOps ML team to ensure telemetry schemas meet feature engineering and ML model requirements.
• Pipeline Operations: Operate the Collector pipeline at scale, managing backpressure handling, sampling strategies (tail/head), and cardinality budgets.
Salary Range- $120,000-$130,000 a year
#LI-OJ1