About the Role:Our Data Platform group at Crowdstrike is unique among its kind for being uncommonly customer-focused. We build and operate systems to centralize all of the data from Falcon Sensors and 3rd Party sources derived from trillions of events/day, and we also drive industry-leading innovation on a hyper scale security data lake that helps find bad actors and stop breaches.
Setting ourselves apart, we make it easy for all customers to utilize the platform for batch and streaming analytics, machine learning and threat hunting through the production and delivery of self-service platforms (Query Platform, Analytics Platform, Enrichment platform...) using Spark and Flink respectively as the runtimes. These self-service platforms allow customers to apply their own custom schema, syntax, data models (etc.) to our historical cyberattack data, modeling threats in a way that empowers them to predictively build (among other things) behavioral automations that defend against such threats 'before' they appear in their own environments.
In this role, you will be a Principal Engineer within Data Platform, owning - in entirely hands-on capacity the design, build and delivery of a new self-service data enrichment platform which will take our customers' proactive defense measures to the next level.
As a leader in data platform, you will contribute to the full spectrum of our systems, including query processing, scalable pipeline builds with largely Apache-based ingestion, materialized view, transformation and data storage frameworks, and tools/applications that make data available to thousands of users and hundreds of internal systems.
What You'll Do:- Shape the vision of our Analytics Data Platform for its next phase of growth: building a Unified Data Catalog as well as Query Analysis for structured data stored in different forms (columnar vs graph), and building and optimizing query performance using different techniques of indexing and data partitioning.
- Design, develop, and maintain a data platform that processes petabytes of data.
- Participate in technical reviews of our products and help us develop new features and enhance stability.
- Continually help us improve the efficiency of our services so that we can delight our customers.
- Help us research, evolve and implement new ways for both internal stakeholders as well as customers to query their data efficiently and extract results in the format they desire.
What You'll Need:- (One among:) 17+ years exp with B.S. in a related field, 15+ years with M.S. in a related field, or 12+ years with PhD in a related field.
- Experience building and supporting very high scale data platform and data storage systems (MINIMUM: 100s of TB/day in either a current or past role)
- Significant experience performance-tuning or developing internals (source code) within Spark, Flink, Iceberg or Pinot or an equivalent structured streaming, Time Series, OLAP or Open Table real time system.
- Production experience building either Spark- or Flink-based self-service data platforms, or equivalent (i.e. with Ray, or building spark- or Flink-like frameworks themselves in Scala, Akk, etc.)
- Strong familiarity with (and ample hands-on experience tuning & optimizing) at least one applicable technology in the Apache Hadoop ecosystem: Spark, Kafka, Hive/Iceberg/Delta Lake, Presto/Trino, Pinot, Druid, etc.
- 3+ years coding in Java, Scala, Kotlin or another JVM language (bonus points for experience tuning the language, i.e. garbage collection, memory management...)
- Production experience with relational SQL and NoSQL databases, including Postgres/MySQL, Cassandra/DynamoDB, etc.
- Proven expertise with multiple big data frameworks in general, especially handling data volume at (ideally) multi-petabyte scale.
- Proven expertise with algorithms, distributed systems design and the software development lifecycle.
- Great test driven development discipline.
- Reasonable proficiency with Linux administration tools.
- Proven experience utilizing AI technologies to enhance decision-making, streamline workflows and processes, improve efficiency and drive business outcomes.
- Proven ability to work effectively with remote teams.
Bonus Points:- Familiarity with Go.
- Familiarity with Kubernetes/Mesos or equivalent.
- Production experience with Flink, especially as runtime for a self-service platform.
#LI-MP2
#LI-SF1
Benefits of Working at CrowdStrike:- Market leader in compensation and equity awards
- Comprehensive physical and mental wellness programs
- Competitive vacation and holidays for recharge
- Paid parental and adoption leaves
- Professional development opportunities for all employees regardless of level or role
- Employee Networks, geographic neighborhood groups, and volunteer opportunities to build connections
- Vibrant office culture with world class amenities
- Great Place to Work Certified™ across the globe
The base salary range for this position for all U.S. candidates is $195,000 - $290,000 per year, with eligibility for bonuses, equity grants and a comprehensive benefits package that includes health insurance, 401k and paid time off.
For detailed information about the U.S. benefits package, please click here.