Staff Engineer - Customer Facing

Data Direct Networks

$150K — $180K *
Technical Services
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years experience in enterprise storage, distributed systems, or cloud infrastructure at a Senior or Staff level.
  • Deep knowledge of file systems and storage technologies such as S3, POSIX, and NFS.
  • Proficient in Linux systems, with expertise in kernel-level troubleshooting.
  • Strong programming skills in Python or C++.
  • Ability to use diagnostic tools like strace, tcpdump, and perf.
  • Desire to engage with customers and take ownership of complex technical issues.

Responsibilities

  • Communicate technical issues to customers and stakeholders clearly.
  • Lead resolution of complex customer escalations and root-cause analysis.
  • Facilitate live incident responses and cross-functional investigations.
  • Diagnose complex issues across system, protocol, and application layers.
  • Document and develop troubleshooting guidance and performance practices.
  • Mentor engineers and influence engineering best practices.
  • Collaborate with Field CTOs and Sales Engineers on strategic issues.
  • Implement AI and automation for improved diagnostics and reliability.

Benefits

  • Participation in an on-call rotation for after-hours support as needed.
Full Job Description
DDN is seeking a Staff Engineer to join our Infinia Core team. This is a hands-on technical role combining deep distributed-systems engineering with direct engagement with customers running Infinia in production.

You'll own complex technical escalations end-to-end, from root-cause analysis and incident response through to mitigation, customer communication and product improvements. You'll also help shape engineering best practice, mentor other engineers and drive the use of AI and automation to improve reliability and diagnostics.

If you love the technical depth but want to stay behind the curtain, this probably isn't the right fit but if you want to combine serious engineering with real customer impact, read on.

What You'll Do
  • Communicate technical issues clearly to customers, engineers and senior stakeholders, including executive audiences.
  • Own complex customer escalations from diagnosis through to resolution, mitigation and RCA.
  • Lead live incident response, war rooms and cross-functional investigations with Engineering, QA and Field teams.
  • Debug complex distributed-systems, storage and performance issues across the system, protocol and application layers.
  • Reproduce customer issues and feed findings into product and reliability improvements.
  • Develop runbooks, troubleshooting guidance and performance-tuning practices.
  • Act as a technical authority on Infinia internals, mentoring engineers and influencing architectural best practice.
  • Partner with Field CTOs, Solutions Architects and Sales Engineers on strategic customer issues.
  • Use AI, automation and observability to improve diagnostics, reliability and MTTR.
  • Communicate technical issues clearly to customers, engineers and senior stakeholders, including executive audiences.
  • This position requires participation in an on-call rotation to provide after-hours support as needed.

What You'll Bring

Must-Haves
  • Significant experience in enterprise storage, distributed systems or cloud infrastructure, with technical leadership at Senior or Staff level.
  • Deep understanding of file systems and storage technologies, including S3, POSIX, NFS and storage performance.
  • Strong Linux systems knowledge, including kernel-level troubleshooting and debugging.
  • Strong coding ability in Python or C++.
  • Proven ability to diagnose complex issues using tools such as strace, tcpdump and perf.
  • Genuine interest in working directly with customers and taking ownership of complex problems through to resolution.

Nice-to-Haves
  • Experience with DDN, VAST, Weka or similar scale-out storage/file systems.
  • Familiarity with observability platforms such as Prometheus, Grafana, ELK or OpenTelemetry.
  • Knowledge of replication, consistency models and data integrity mechanisms.
  • Experience supporting AI/ML, LLM training or other high-performance computing environments.
  • Experience using AI tools for log analysis, troubleshooting, automated RCA or reducing MTTR.

Similar Jobs

More Jobs at Data Direct Networks

More Technical Services Jobs

Find similar Staff Engineer - Customer Facing jobs: