Bachelor's degree in Computer Science, Information Systems, or equivalent experience
2+ years in a hands-on data development role with production pipelines
Experience owning a data pipeline component in a production environment
Strong SQL skills including complex queries and performance diagnosis
Proficiency in production-grade Python for ETL and tooling
Hands-on experience with Spark on EMR or similar technologies
Proficient with Airflow for orchestration and job management
Experience with data ingestion tools in a production setting
Familiarity with OLAP engines like StarRocks or ClickHouse
Solid understanding of data warehouse modeling and concepts
Comfortable working with AWS services and tools
Responsibilities
Own a data platform component end to end, ensuring quality and performance
Design and build ETL processes using EMR (Spark) and Airflow
Model data for analytical use cases, applying dimensional modeling techniques
Implement and monitor data quality checks and troubleshoot issues
Optimize job performance and infrastructure costs based on workload
Develop reusable components and automate repetitive tasks
Collaborate with analysts and stakeholders to transform requirements into actionable models and datasets
Benefits
Free snacks and drinks
Fully paid medical, dental, and vision insurance (partial coverage for dependents)
Contributions to 401K funds
Bi-annual reviews with potential for annual pay increases
Health and wellness benefits, including a free gym membership
Quarterly team-building events
Full Job Description
KEY RESPONSIBILITIES
Own a component of the data platform end to end - its pipelines, tables, quality checks, monitoring, and recovery.
Design and build ETL on EMR (Spark) and Airflow, and data ingestion from OLTP sources via DataX / CDC, including incremental sync and upstream schema changes.
Model the data for your area: dimensional models and layering (ODS / DWD / DWS / ADS) that serve real analytical use cases.
Own data quality for your component - write the checks and freshness monitoring, and resolve issues before they reach downstream users.
Tune runtime and infrastructure cost for the jobs you own, and choose the right engine for a workload (StarRocks serving vs. batch on EMR).
Build components other engineers can reuse, and automate work that would otherwise repeat.
Work directly with analysts and business stakeholders to turn requirements into models and datasets that actually get used.
Requirements
REQUIRED QUALIFICATIONS
Bachelor's degree in Computer Science, Information Systems, or a related field, or equivalent practical experience.
2+ years of hands-on data development in a production environment - scheduled pipelines with real downstream consumers.
Experience owning a pipeline or component in production: you were the person accountable for it, including when it broke.
Strong SQL: window functions, complex multi-table joins, incremental and idempotent writes; able to read an execution plan and diagnose data skew or slow queries.
Production-grade Python: maintainable, testable ETL and tooling code (PySpark / pandas / boto3) with proper error handling and logging.
Hands-on with Spark on EMR, or equivalent Hadoop / Hive experience, including partitioning, shuffle, and memory tuning.
Production experience with Airflow: DAG design, dependencies, retries, idempotent reruns, backfills, and SLA alerting.
Experience with a data ingestion or CDC tool in production (DataX, Sqoop, Debezium, or similar): extracting from OLTP sources, incremental sync, and handling upstream schema changes.
Experience with an OLAP engine - StarRocks, Doris, or ClickHouse: table models, partitioning and bucketing, materialized views, and query tuning.
Solid grasp of data warehouse modeling: dimensional modeling, slowly changing dimensions, and layering conventions.
Comfortable on AWS (S3 with Parquet / ORC and sensible partitioning, IAM basics, day-to-day EMR operations), Linux, and Git.
Comfortable picking up new tools as the platform evolves - we'd rather hire someone who learns a new engine quickly than someone who has only ever used ours.
Effective use of AI to solve data problems - using it to move faster on SQL, debugging, and unfamiliar schemas, with the judgment to catch output that looks right but isn't.
PREFERRED QUALIFICATIONS
AWS cost optimization: EMR instance sizing and Spot strategy, S3 lifecycle policies, StarRocks vs. Athena trade-offs.
Lakehouse table formats: Iceberg, Hudi, or Delta.
Streaming: Kafka with Flink or Spark Structured Streaming.
dbt, Glue Data Catalog, data lineage, or a data quality framework.
QuickSight dataset, SPICE, and row-level permission design.
Benefits
Salary range: TBD
Free snacks and drinks
Fully paid medical, dental, and vision insurance (partial coverage for dependents)
Contributions to 401K funds
Bi-annual reviews, and annual pay increases
Health and wellness benefits, including free gym membership