Datadog

Senior Software Engineer - Incident Insights & Readiness

Datadog • $192K — $240K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years of software development experience, particularly in Go and Python.
  • Experience with distributed systems and Kubernetes, understanding complex failure modes.
  • Proven ability to manage ambiguous technical challenges from design to delivery.
  • Skilled in analyzing incidents and driving engineering improvements based on operational insights.
  • Experience in on-call rotations and enhancing incident response processes, with incident commander experience as a plus.
  • Strong empathy, collaboration, and communication skills to build relationships across teams.
  • Experience mentoring engineers and influencing technical direction without formal authority.

Responsibilities

  • Own and enhance the on-call experience by establishing best practices and building supportive platforms.
  • Lead the design and implementation of software to streamline incident response processes.
  • Collaborate on post-mortem processes to identify learning opportunities and reduce friction.
  • Facilitate incident reviews that promote learning and a blameless culture across teams.
  • Provide technical leadership and coaching to team members to foster their growth.
  • Train on-call engineers in incident management best practices and processes.
  • Drive cross-functional initiatives to improve reliability and operational excellence across engineering teams.

Benefits

  • New hire stock equity (RSUs) and employee stock purchase plan (ESPP).
  • Continuous professional development and career pathing opportunities.
  • Intradepartmental mentor and buddy program for networking.
  • Inclusive company culture with Community Guilds for employee resource groups.
  • Access to internal panel discussions on inclusion.
  • Free global mental health benefits for employees and dependents age 6+.
  • Competitive global benefits package.
Full Job Description
The Incident Insights & Readiness SRE team at Datadog fosters a resilient culture by using incidents as learning opportunities and catalysts for growth. Our users are Datadog engineers, and we build the software, tooling, and operational frameworks that help them prepare for, respond to, and learn from incidents. We work closely with engineering teams across Datadog to analyze incidents and turn those insights into better tools, stronger incident response, and organizational learning. Our efforts empower Datadog to navigate unexpected failures confidently, efficiently, and with a commitment to continuous learning and systems improvement.

At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them.

What You'll Do:
  • Own and improve the on-call experience for the company by establishing best practices and building platforms to support on-call rotations and compensation.
  • Define how we respond to incidents, lead the design and implementation of software to streamline the process, and collaborate with product teams to improve incident response across Datadog. Our aim is to fully support our incident responders in dealing with complexity.
  • Contribute to the post-mortem process for the company, collaborating with teams on writing them, and identifying opportunities to reduce friction and enhance learning value for the organization. Our team also runs a weekly postmortem reading group.
  • Support various teams in facilitating incident reviews that emphasize learning and blamelessness. Help them share their learnings across the organization to improve the resilience of our people.
  • Provide technical leadership and day-to-day coaching to team members, accelerating their growth through design reviews, collaborative problem-solving and operational excellence best practices.
  • Train our on-callers in incident and post-mortem processes, sharing expertise in incident management best practices. This involves both introducing newcomers to on-call responsibilities and refreshing the knowledge of existing engineers.
  • Lead cross-functional initiatives in engineering organizations across Datadog, embedding with teams to understand their challenges and drive lasting improvements to reliability and operational excellence.


Who You Are:
  • At least 5 years of experience building software that solves real user problems. Experience designing new features and collaborating on code and technical design reviews. We primarily develop in Go and Python, with a bit of TypeScript.
  • Experience building or operating distributed systems, with familiarity with Kubernetes and an understanding of complex failure modes.
  • Demonstrated ability to independently own ambiguous technical problems from design through delivery while balancing long-term engineering quality with pragmatic execution.
  • Experience analyzing incidents, identifying systemic risks, and driving engineering improvements informed by operational learnings.
  • Experience participating in on-call rotations and improving incident response processes. Experience serving as an incident commander or incident coordinator is a plus.
  • Empathy, collaboration, and communication skills in English to cultivate strong relationships across various teams in the organization
  • Experience mentoring engineers, driving cross-functional initiatives, and influencing technical direction without relying on organizational authority.
  • We welcome candidates from a variety of backgrounds, including software engineering, site reliability engineering, production engineering, infrastructure, and other roles focused on building reliable systems or improving incident response.


Datadog values people from all walks of life. We understand not everyone will meet all the above qualifications on day one. That's okay. If you're passionate about technology and want to grow your skills, we encourage you to apply.

Benefits and Growth:
  • New hire stock equity (RSUs) and employee stock purchase plan (ESPP)
  • Continuous professional development, product training, and career pathing
  • Intradepartmental mentor and buddy program for in-house networking
  • An inclusive company culture, ability to join our Community Guilds (Datadog employee resource groups)
  • Access to Inclusion Talks, our internal panel discussions
  • Free, global mental health benefits for employees and dependents age 6+
  • Competitive global benefits


Benefits and Growth listed above may vary based on the country of your employment and the nature of your employment with Datadog.

Datadog offers a competitive salary and equity package, and may include variable compensation. Actual compensation is based on factors such as the candidate's skills, qualifications, and experience. In addition, Datadog offers a wide range of best in class, comprehensive and inclusive employee benefits for this role including healthcare, dental, parental planning, and mental health benefits, a 401(k) plan and match, paid time off, fitness reimbursements, and a discounted employee stock purchase plan.

The reasonably estimated yearly salary for this role at Datadog is:

$192,000-$240,000 USD

About Datadog

Datadog is a monitoring and analytics platform for cloud-scale infrastructure and applications. The company was founded in 2010 by Olivier Pomel and Alexis Lê-Quôc and is headquartered in New York City. Datadog's platform assists organizations in improving agility, increasing efficiency, and providing end-to-end visibility across dynamic or high-scale infrastructures. The company's SaaS-based data analytics platform integrates and automates infrastructure monitoring, application performance monitoring, and log management to provide unified, real-time observability of customers' entire technology stack. Datadog's customers include Airbnb, Twilio, and The Washington Post.
Learn more about Datadog
Size
3,200 employees
Market Cap
$22.5 billion
Industry
Net Income
-$24.5 million
Founded
2010
Revenue
$603.4 million
NASDAQ

Similar Jobs

More Jobs at Datadog

More Information Technology Jobs

Find similar Senior Software Engineer - Incident Insights & Readiness jobs: