Sr. Site Reliability Engineer - Core Platform & Embedded Reliability (Hybrid)

CrowdStrike Holdings, Inc.$140K — $215K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 10+ years of experience building and operating large-scale distributed systems
  • 5+ years of developing microservices in modern backend languages like Go or Python
  • Expert-level proficiency in at least one programming language, preferably Go
  • Deep understanding of distributed systems and their failure modes
  • Proven experience in scaling backend systems and performance optimization
  • Strong systems thinking and ability to influence across organizational boundaries
  • Degree in Computer Science or equivalent experience

Responsibilities

  • Partner with engineering leadership to drive multi-year reliability roadmaps
  • Design and implement architectural improvements impacting all CrowdStrike teams
  • Develop services to meet aggressive reliability and scalability demands
  • Extend and build libraries for cross-cutting concerns in the cloud platform
  • Lead initiatives around reliability, scalability, and cost efficiency
  • Establish observability practices to drive automation and improve performance
  • Conduct resilience engineering and design automation to improve reliability

Benefits

  • Comprehensive physical and mental wellness programs
  • Competitive vacation and holidays
  • Paid parental and adoption leaves
  • Professional development opportunities for all employees
  • Employee Networks and volunteer opportunities to build connections
  • Vibrant office culture with world-class amenities
  • Great Place to Work Certified™ globally
Full Job Description
About the Role:

CrowdStrike Falcon is the industry standard in cloud-native cybersecurity and threat hunting, processing trillions of events per day. As a Principal SRE, you will operate at the intersection of our Core Platform and Embedded Reliability charters: building the foundational libraries, services, and tooling that every product group depends on, while embedding directly with product engineering teams and their leadership to drive reliability outcomes at scale.

While we embrace the SRE moniker, at CrowdStrike it means something far more service-oriented and engineering-heavy than traditional operations. This is hands-on systems engineering - writing production code, re-architecting critical systems, and eliminating entire classes of failure - not ticket management. It is far and away our most self-driven and autonomous backend engineering role, with the freedom to move up, down, and laterally across the stack as needed.

You'll join a group with bottom-up visibility and ownership of a fast-expanding codebase (including AI-based Detection & Response) and a far-reaching mandate for service resiliency. Several key feature additions and critical re-architecture initiatives are underway: deploying core services to additional cloud providers, modularizing into reusable components, improving core libraries and frameworks, maturing observability tooling (tracing, profiling, alerting, SLOs), and automating away manual toil across all of the above. Recent examples of the team's work include introducing adaptive concurrency into our core Kafka library to implement flow control that protects database performance, resolving critical issues in leader election libraries, and building infrastructure-as-code tooling that eliminated manual deployment processes.

At the Senior Engineer level, your influence is organizational. Product engineers and engineering leaders will come to you for guidance on architectural decisions because you've earned credibility through hands-on work and delivered results. You will shape architectural choices that affect every feature development team and provide shared architectural components leveraged throughout the Falcon Platform - including Unified Search, Protobuf libraries, and other shared-tier infrastructure.

Why This Role Matters: Our customers depend on us to protect their businesses from sophisticated threats, and reliability isn't optional - it's fundamental to our mission. Your work directly impacts whether organizations around the world can defend themselves against cyberattacks. You'll work on problems that matter, at a scale few companies can match, with the autonomy to make real architectural decisions.

Location: This position requires candidates to be based in Midtown Manhattan, NY; Redmond, WA; Sunnyvale, CA; or Austin, TX. It's a hybrid role, with employees typically working 2 to 3 days per week in the office.

What You'll Do:
  • Partner with engineering leadership across multiple product groups to define and drive multi-year reliability roadmaps.
  • Design and implement architectural improvements to services, libraries, and platforms that impact teams across all of CrowdStrike.
  • Develop and maintain services that meet aggressive reliability and scalability demands.
  • Extend and build new libraries for cross-cutting concerns spanning CrowdStrike's cloud platform, which comprises hundreds of libraries and services.
  • Lead initiatives around reliability, scalability, performance, and cost efficiency in large-scale distributed systems.
  • Establish foundational observability practices: ensure teams instrument services properly, react to signals effectively, and leverage observability to drive automation such as continuous delivery.
  • Define and implement service-level objectives and error budgets that drive real decision-making and prioritization.
  • Lead performance and cost optimization efforts: profiling, bottleneck analysis, capacity planning, and cloud efficiency improvements.
  • Conduct resilience engineering: chaos experiments, failure injection, failure modeling, and designing for graceful degradation.
  • Design and implement automation and infrastructure-as-code to improve infrastructure reliability and eliminate manual toil.
  • Provide technical leadership during complex incidents and ensure follow-through on retrospectives with concrete improvements that eliminate entire classes of failures.
  • Identify opportunities to extract common patterns into shared libraries and tools, or partner with platform teams on improvements benefiting multiple product groups.
  • Continuously re-evaluate our products to improve architecture, knowledge models, developer and user experience, performance, and stability.
  • Drive strategic technical decisions and influence infrastructure and operational improvements across the organization.
  • Mentor and coach engineers, raising the technical IQ of the team and driving architectural standards across the org.
  • Use and give back to the open source community; evangelize software engineering best practices, especially as they pertain to Go.
  • Brainstorm, define, and build collaboratively with members across multiple teams as an energetic self-starter who owns and is accountable for deliverables.


What You'll Need:
  • 10+ years of experience building and operating distributed systems and service-oriented backends at scale.
  • 5+ years developing microservices for a SaaS product in a modern backend language (Go, Java, Scala, Kotlin, Python, Node.js).
  • Expert-level proficiency in at least one programming language, with expert-level Go or demonstrated ability and willingness to reach expert level in Go.
  • Deep understanding of distributed systems: consensus algorithms, replication, consistency models, failure modes, and scalability patterns.
  • Proven experience scaling backend systems - sharding, partitioning, horizontal scaling, capacity planning, and performance optimization are second nature.
  • Deep understanding of multi-threading, concurrency, and parallel processing.
  • Track record of making impactful architectural decisions at organizational scope and seeing them through to production.
  • Strong systems thinking and the ability to influence without direct authority across organizational boundaries.
  • Thorough command of engineering best practices: appropriate testing paradigms, effective peer code review, and resilient architecture.
  • Ability to thrive in a fast-paced, test-driven, collaborative, and iterative environment; strong team-player orientation.
  • A desire to ship code and a love of seeing your bits run in production.
  • Degree in Computer Science, or commensurate experience in data structures, algorithms, and distributed systems.
  • Proven experience utilizing AI technologies to enhance decision-making, streamline workflows and processes, improve efficiency and drive business outcomes.


Bonus Points:
  • Experience driving reliability improvements in organizations with hundreds or thousands of microservices.
  • Deep knowledge of Kubernetes or other large-scale orchestration systems.
  • Hands-on experience with AWS, Cassandra, Kafka, Elasticsearch/OpenSearch, or similar large-scale distributed technologies.
  • Experience with Google Cloud Platform (GCP).
  • Experience with Oracle Cloud Infrastructure (OCI).
  • Experience delivering or operating services across multiple cloud providers, including multi-cloud abstraction layers, portability, and cloud-agnostic tooling.
  • Track record of building internal platforms, developer platforms, or tools that other engineers depend on.
  • Experience with infrastructure cost optimization at scale.
  • Background in performance engineering: profiling, optimization, and identifying system bottlenecks.
  • Experience with chaos engineering or resilience testing practices.
  • History of establishing SLO/SLI frameworks and error budgets in production environments.
  • Contributions to the open source community (GitHub, Stack Overflow, technical blogging).
  • Prior experience in the cybersecurity or intelligence fields.


#LI-SC1

Benefits of Working at CrowdStrike:
  • Market leader in compensation and equity awards
  • Comprehensive physical and mental wellness programs
  • Competitive vacation and holidays for recharge
  • Paid parental and adoption leaves
  • Professional development opportunities for all employees regardless of level or role
  • Employee Networks, geographic neighborhood groups, and volunteer opportunities to build connections
  • Vibrant office culture with world class amenities
  • Great Place to Work Certified™ across the globe


The base salary range for this position for all U.S. candidates is $140,000 - $215,000 per year, with eligibility for bonuses, equity grants and a comprehensive benefits package that includes health insurance, 401k and paid time off.

For detailed information about the U.S. benefits package, please click here.

About CrowdStrike Holdings, Inc.

CrowdStrike Holdings, Inc. Careers

Joining CrowdStrike Holdings, Inc. presents an unparalleled opportunity to advance a career in the tech industry with a company at the forefront of digital security. As a leader in cybersecurity solutions, CrowdStrike Holdings, Inc. offers a range of job opportunities that cater to a variety of skills and experiences, from entry-level positions to senior leadership roles.

Explore Job Opportunities

CrowdStrike Holdings, Inc. is continuously seeking talented individuals who are passionate about protecting organizations against cyber threats. With a commitment to innovation and excellence, the company is hiring professionals who are eager to contribute to a team that values hard work and creative solutions.

Innovation and Professional Growth

At CrowdStrike Holdings, Inc., employees are encouraged to push the boundaries of technology and leadership. The company supports professional growth through robust training programs, including leadership development and diversity training, ensuring that every team member has the resources to thrive in their career.

Culture and Benefits

The culture at CrowdStrike Holdings, Inc. is dynamic and inclusive, fostering a workplace where diversity is celebrated and every voice is heard. Employees enjoy comprehensive benefits that support both their professional and personal lives, enhancing job satisfaction and team morale.

Internship Programs

For those starting their career, CrowdStrike Holdings, Inc. offers internship programs that provide a rich learning environment. Interns gain hands-on experience, working alongside seasoned professionals and participating in projects that deliver real-world solutions.

Networking and Career Advancement

CrowdStrike Holdings, Inc. emphasizes the importance of networking within the industry, offering numerous opportunities for employees to connect with thought leaders and innovators. These connections can lead to career advancement and a deeper understanding of the cybersecurity landscape.

Applying for a Position

To apply for a position at CrowdStrike Holdings, Inc., candidates should prepare a resume that highlights relevant experience and skills. The interview process is designed to assess not only professional qualifications but also a candidate's fit within the company culture and team.

Stay Connected with CrowdStrike Careers

Interested candidates can stay informed about new openings and company news by subscribing to job alert emails. This personalized service ensures that potential applicants are the first to know about new opportunities that match their career interests and skills.

Join the Team

CrowdStrike Holdings, Inc. is looking for curious, creative, and solution-driven team players. Explore the employment opportunities on the CrowdStrike Holdings, Inc. careers page to find a position that matches your skills and passions.

SEARCH CROWDSTRIKE JOBS

Keep Up to Date

Stay ahead with career tips, insider perspectives, and industry-leading insights you can put to use today—all from the professionals who work at CrowdStrike Holdings, Inc.

READ CAREERS BLOG

Job Alert Emails

Customize your subscription to receive job alerts, latest news, and insider tips tailored to your preferences. Discover the exciting and rewarding career opportunities waiting at CrowdStrike Holdings, Inc.
Learn more about CrowdStrike Holdings, Inc.

Similar Jobs

More Jobs at CrowdStrike Holdings, Inc.

More Information Technology Jobs

Find similar Sr. Site Reliability Engineer - Core Platform & Embedded Reliability (Hybrid) jobs: