We are looking for a hands-on
Staff Software Engineer who will help design and build our Cluster management platform.
What will you bring to DDN- 8+ years of backend development experience (with a target of matching our senior engineering standards), including deep proficiency in Go for building high-performance, low-overhead system components.
- Production-Scale Observability Expertise: Hands-on experience implementing and operating telemetry pipelines for software-defined clustered, distributed, or cloud-native solutions.
- Deep Mastery of the Metrics Stack: Strong technical knowledge of Prometheus (operators, alerting rules, scraping mechanics) and the VictoriaMetrics stack for long-term, high-cardinality storage.
- Telemetry Industry Standards: Practical experience with the OpenTelemetry (OTel) ecosystem, including custom OTel collector configurations, instrumentation SDKs, and data processing.
- Systems-Level Troubleshooting: A solid understanding of Linux networking, filesystems, and how clustered storage applications behave under heavy I/O workloads.
- Collaborative Mindset: Proven ability to work effectively across geographically distributed teams, driving technical clarity through code reviews and clear documentation.
What you have achieved...- Built and Maintained Telemetry Pipelines: A proven track record of developing or extending proprietary Go components to efficiently ingest, process, and forward massive streams of metrics, logs, and traces.
- Optimized Resource Consumption: Experience managing the CPU and memory footprint of monitoring agents to ensure they do not compete with core storage data paths.
- Strong Testing & Regression Habits: Dedication to writing robust unit and integration tests to ensure telemetry components remain stable during live cluster upgrades.
- Independent Feature Delivery: A proven ability to take ownership of complex technical initiatives and independently make progress in a fast-paced environment.
- Concept Visualization: Ability to turn abstract cluster state data into logical, well-structured telemetry frameworks that bring visibility to complex system scenarios.
What will you be doing- Execute the Telemetry Architecture: Take on core projects within the observability domain, ensuring seamless integration between our proprietary Go infrastructure and open-source tools.
- Optimize the Observability Stack: Help design and refine how OpenTelemetry, Prometheus, and VictoriaMetrics handle the massive metrics volume generated by our storage cluster.
- Drive Code Excellence: Act as a key technical contributor to our Software Defined Storage control plane, writing clean, performant Go code and providing rigorous code reviews.
- Full Lifecycle Engineering: Participate actively within the Scrum model-from initial design and coding to automated testing, usability reviews, and release.
- Document and Standardize: Ensure our telemetry frameworks are well-documented, making it easy for other engineering teams to instrument their components properly.
- Global Support Rotation: Contribute to our global team on-call rotation, leveraging your own observability tools to provide high-level technical support for our distributed footprint.