Research Engineer, Content Understanding

Exa Labs Inc.

• $135K — $160K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Graduate-level ML experience (Master's or PhD with at least 2 years of relevant experience), or an exceptionally strong undergrad
  • Ability to build a transformer from scratch in PyTorch and train models for cost-effective deployment
  • Experience in building large-scale datasets and focusing on data supervision
  • Comfortable with undefined ground truths and defining them as part of the role
  • Passionate about improving knowledge quality and its global impact

Responsibilities

  • Develop and enhance parsing capabilities for web pages
  • Create models to assess and standardize page quality
  • Address credibility and misinformation as a modeling challenge
  • Determine semantic similarities between documents for effective deduplication
  • Build scalable classification and extraction systems
  • Design supervision methods for unlabeled data
  • Trace and rectify poor search results back to their source

Benefits

  • Flexible project assignments based on skills and interests
  • Opportunity to work on cutting-edge search architecture
  • Engagement with unsolved problems in AI and document understanding
  • Collaborative environment focused on high-quality knowledge extraction
  • Potential for significant impact on global information quality
Full Job Description
As a backend engineer, youd play a critical role in our search architecture. Were pretty flexible on what projects people work on based on their skills and interests.

Search quality is bounded by what we understand about a page. Before anything can be retrieved, something has to work out what the page actually says. That means parsing it into the parts that are content and the parts that are furniture, classifying what kind of page it is and what it is about, telling whether the page is usable at all, extracting when it was published, judging how good it is and whether it can be trusted, and working out whether it says anything that a page we already have does not. All of this has to work on every page on the web, in every language, in every shape the web comes in.

Some of this is classic document understanding. Some of it is much more open. Credibility and misinformation, AI-generated and machine-spun content, and pages written to be found rather than read are all unsolved, and search results are only as trustworthy as our answers to them.

We are looking for a research engineer to work on this. There is a lot of room to do it well.

Desired Experience
  • Graduate-level ML experience (Masters or PhD with at least 2 years of relevant experience), or an exceptionally strong undergrad
  • You can build a transformer from scratch in PyTorch, and you have trained models that then had to be cheap enough to run everywhere
  • You like building large-scale datasets and living in the data. Most of the wins here are in the supervision rather than the architecture
  • You are comfortable with problems where the ground truth does not exist yet and defining it is part of the job
  • You care about the problem of finding high quality knowledge and recognize how important this is for the world


Example Projects
  • Make parsing work on the pages where it currently does not, and prove the improvement rather than assert it
  • Teach a model to judge page quality, and get everyone to agree on what quality means well enough to supervise it
  • Work on credibility and misinformation as a modelling problem: what a page claims, whether it is a reliable source of it, and whether it was written for a reader or for a crawler
  • Decide whether two documents are semantically the same or genuinely different, so we can deduplicate the web without collapsing pages that a user would want to see separately
  • Build classification and extraction that is accurate at web scale and cheap enough to run on all of it
  • Design the supervision for something nobody has labels for, and find out whether it is learnable at all
  • Trace a bad search result back to the page-level prediction that caused it, and fix it at the source


Similar Jobs

More Jobs at Exa Labs Inc.

More Information Technology Jobs

Find similar Research Engineer, Content Understanding jobs: