Job Description: We are seeking machine learning engineers to join our team full-time. As part of your role, you will help us build pipelines of data collection, data extraction, data filtering/synthetic data generation and data analysis. You will own all work related to acquiring high-quality data to power the training of our domain-specific models end to end. You will work closely with other researchers and engineers to empower our next generation of domain-specific models. We value rapid prototyping, iterating, and shipping new systems quickly.
Required Qualifications: - BS/MS/PhD in Computer Science or a related field.
- Proficiency in at least one deep learning framework, such as PyTorch.
- Experience in machine learning projects in text or vision, e.g., has trained machine learning models to tackle a specific problem.
- Strong expertise in large stateful distributed systems and data processing.
- Strong proficiency in building large-scale data processing pipelines, familiar with distributed workload (e.g., multiprocessing, Ray, Docker, Kubernetes).
- Proficiency in at least one programming language commonly used in machine learning, such as Python and ability to write clean, maintainable code.
- Excellent problem-solving skills and attention to detail, especially when handling data anomalies and biases to further improve data quality.
Key Competencies - Active Github contributions are a big plus.
- Experience in building large-scale datasets.
- Familiar with at least one of the following tools for data crawling (e.g. Scrapy), data collection (e.g., VPNs, Selenium), data processing (e.g., Hadoop, Datasketch).
- Building bespoke data processing libraries from scratch.
- Keeping up with state-of-the-art techniques for preparing AI training data.
- Organizing and meticulously bookkeeping data across multiple clouds, of multiple modalities, and from many sources.
- Multilingual which contributes to enriching the language diversity crucial for robust model training.
Responsibilities: - Design and develop data processing pipelines, including data extraction, data filtering, data labeling, etc.
- Implement machine learning models to improve the quality and diversity of data (especially in the data extraction stage), e.g., quality classifier, document layout model, code verification model, etc.
- Own and lead engineering projects in the area of data acquisition, including web crawling, data ingestion, and processing.
- Collaborate with our Applied Research, Technology, and Architecture teams to ensure smooth data flow and system operability.
- Develop and deploy highly scalable distributed systems capable of handling terrabytes of data.
- Architect and implement algorithms for data indexing and search capabilities.
- Build and maintain backend services for data storage, including work with key-value databases and synchronization.
- Deploy solutions in a Kubernetes Infrastructure-as-Code environment and perform routine system checks.