About the RoleYou'll design, build, and maintain the robust data pipelines and infrastructure required for large-scale AI research. Your work will ensure that researchers and engineers have clean, reliable, and high-quality data for every experiment.
Responsibilities- Build and operate ETL pipelines for large, heterogeneous datasets.
- Develop and manage storage, cleaning, labeling, and data access infrastructure.
- Build tools for data quality, validation, and monitoring.
- Collaborate to create datasets and benchmarks for novel ML problems.
Requirements- Strong software engineering skills (Python, SQL, PySpark, data infra tools).
- Experience scaling data pipelines in cloud or hybrid environments.
- Ability to work with unstructured, messy, or adversarial data.
- Curiosity and rigor about data quality, privacy, and reproducibility.