Benchling

Software Engineer, Model Evaluation and Improvement

Benchling$130K — $155K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 2+ years of experience at the intersection of biology and AI, particularly in evaluating scientific models or LLMs.
  • Experience with LLMs, understanding their strengths and weaknesses in biological contexts.
  • Strong curiosity and excitement about frontier AI, with a desire to explore and expand model capabilities.
  • Ability to tackle ambiguous problems in a fast-evolving technical landscape.
  • Collaborative mindset for engaging with engineers, scientists, and external partners.
  • Familiarity with fast-paced work environments where priorities can shift rapidly.

Responsibilities

  • Build datasets for evaluating and improving frontier models from complex scientific data.
  • Analyze model failure modes to identify improvement opportunities in frontier AI.
  • Build scalable data infrastructure for curating and validating scientific datasets.
  • Collaborate with AI labs to develop new approaches for scientific model evaluation.
  • Translate expert scientific judgment into clear problems and evaluation criteria.

Benefits

  • In-person collaboration in a dynamic office environment in San Francisco.
  • Encouragement of rapid experimentation and innovation.
  • Exposure to cutting-edge work in AI and biology fields.
Full Job Description
Role Overview

We're a team focused on making frontier AI models better at science. LLMs know an extraordinary amount of biology, but there's still a large gap in reasoning for the real-world problems scientists face every day. We recently published some of our work here.

You'll build the datasets, evaluations, and systems that help close that gap. You'll work with scientists to turn complex scientific work into rigorous tasks that models can learn from and be evaluated against. You'll partner with leading AI labs to understand where models fail and how to improve them.

This is an early and rapidly evolving area. You'll work at the intersection of software engineering, biology, and frontier AI: finding tasks that are challenging for LLMs and valuable to scientists, designing evaluations that capture real scientific judgment, and building systems to create these tasks at scale.

RESPONSIBILITIES
  • Build datasets for evaluating and improving frontier models, turning complex scientific data into high-quality tasks and environments for LLMs.
  • Analyze model failure modes, running experiments across frontier models to understand where they struggle and identify opportunities for improvement.
  • Build scalable data infrastructure, creating pipelines that curate, transform, and validate large volumes of scientific data into tasks for model evaluation and improvement.
  • Collaborate with frontier AI labs, helping develop and evaluate new approaches for improving models on challenging scientific tasks.
  • Work closely with scientists, translating expert judgment into problems and evaluation criteria that can reliably distinguish strong model behavior.
QUALIFICATIONS
  • 2+ years at the intersection of biology and AI, with experience evaluating and improving scientific models or LLMs for biological applications.
  • Experience building with LLMs, with an intuition for where current models excel, where they struggle, and how to design systems around their capabilities.
  • Curiosity and excitement about frontier AI, with a desire to understand and push the capabilities of rapidly improving models.
  • Comfort working on ambiguous problems, where the playing field is rapidly shifting and the right technical approaches are still being discovered.
  • Collaborative mindset, able to work closely with engineers, scientists, and external research partners.
  • Desire to work in a fast-paced environment, where priorities can shift and rapid experimentation is encouraged.


HOW WE WORK

This is an in-person team in San Francisco built around collaborating in the office in a fast-paced environment. We're in the office Monday through Friday.

#LI-KW1

About Benchling

Benchling is a cloud-based informatics platform that accelerates life sciences R&D by streamlining workflows and centralizing data. The platform offers a suite of applications for molecular biology, including DNA design, antibody design, CRISPR analysis, and protein expression. Benchling's customers include pharmaceutical companies, biotechs, and academic institutions. The company was founded in 2012 and is headquartered in South San Francisco, California.
Learn more about Benchling
Size
500 employees
Industry
Founded
2012

Similar Jobs

More Jobs at Benchling

More Information Technology Jobs

Find similar Software Engineer, Model Evaluation and Improvement jobs: