Web Crawling - Research Engineer

Thinking Machines Lab

• $350K — $475K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 8+ years in designing and scaling web crawlers or large data-acquisition systems
  • Proven ownership of crawler infrastructure at internet scale
  • Proficient in Python, Go, or Rust with experience in distributed systems
  • Familiar with legal aspects of large-scale web data collection like robots.txt and rate limiting

Responsibilities

  • Design and scale the web crawler for Inkling's pretraining data
  • Build pipelines for data extraction and quality filtering
  • Create specialized crawlers for high-value data
  • Collaborate with pretraining team on data impact on model performance
  • Enhance reliability and efficiency of ingestion systems at petabyte scale
  • Set technical direction and mentor engineers in crawling best practices

Benefits

  • Generous health, dental, and vision benefits
  • Unlimited PTO
  • Paid parental leave
  • Relocation support as needed
Full Job Description
About the Role

We're hiring a Software Engineer to build and own our web-crawling systems, from distributed collection at internet scale through filtering, deduplication, and deciding what data we keep.

The ideal candidate has built and scaled a web crawler or large-scale data-acquisition systems. In this role, you'll write and own production systems: the crawler itself, the infrastructure that runs it at scale, and the pipelines that turn raw crawls into usable pretraining data. You'll work closely with our pretraining and data teams to understand what's actually moving model quality, but this is fundamentally an engineering role, not a research one.

What You'll Do
  • Design and scale the web crawler and ingestion infrastructure that sources Inkling's pretraining data
  • Build pipelines for large-scale extraction, deduplication, and data quality filtering
  • Build specialized crawlers for high-value or hard-to-reach data sources
  • Work with the pretraining team to understand how changes in crawled data affect model performance
  • Improve the reliability and efficiency of crawling and ingestion infrastructure at petabyte scale
  • Help set technical direction for this area as it grows, and bring other engineers up to speed on what you've learned


Skills & Qualifications
Minimum Qualifications
  • 8+ years designing, building, and scaling web crawlers, scrapers, or large-scale distributed data-acquisition systems
  • A track record of owning crawler or data-acquisition infrastructure at internet scale
  • Strong software engineering skills in a language such as Python, Go, or Rust, with real experience in distributed systems
  • Working knowledge of the practical and legal considerations of large-scale web data collection (robots.txt, rate limiting, licensing)
Preferred Qualifications
  • Experience applying machine learning to crawl selection, extraction, or data quality classification at internet scale
  • Experience setting technical direction for a crawling, data acquisition, or search infrastructure team, whether or not that was your formal title
  • Experience designing systems for petabyte-scale storage and processing
  • Track record of open-source contributions to crawling, scraping, or data infrastructure tools
  • Background at a search engine (crawling, indexing) or a frontier AI lab's data acquisition team


Logistics
  • Location: This role is based in San Francisco, CA.
  • Compensation: Depending on background, skills and experience, the expected annual salary range for this position is $350,000-$475,000 USD.
  • Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.
  • Benefits: Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.

Similar Jobs

More Jobs at Thinking Machines Lab

More Information Technology Jobs

Find similar Web Crawling - Research Engineer jobs: