Member of Technical Staff, Frontier Evals

Arcada Labs

$150K — $180K *
Consumer Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Strong STEM background in areas like Computer Science, Data Science, Statistics, Math, Engineering, or Physics.
  • Deep curiosity about frontier AI models and their performance.
  • Eager to explore model behavior and areas of failure.
  • Intellectual drive to engage with cutting-edge AI development.
  • Willingness to engage in hands-on work and problem solving.

Responsibilities

  • Define measurements for frontier AI models.
  • Design benchmarks and evaluation metrics that set industry standards.
  • Conduct experiments and analyze model performance under real-world conditions.
  • Investigate failures in models to uncover insights about capabilities.
  • Publish research and technical documentation influencing AI evaluation practices.

Benefits

  • Location in San Francisco with visa sponsorship and relocation support.
  • Unique Sunday-Friday work schedule, providing Saturday off.
  • Opportunity to work on industry-leading projects followed by influential figures.
  • Competitive salary plus meaningful equity ownership during critical growth phase.
Full Job Description
Role

You'll define how frontier AI models are measured. You'll design new benchmarks, run experiments, analyze model behavior, and build evaluation methodologies that become trusted signals for the industry. Your work will shape our public leaderboards and the evaluation tools we share with frontier labs.

Here's an example of a piece of industry-leading work done in this field. This is a SOTA STS benchmark, advised by OpenAI: https://audioarena.ai/. Email us for the pre-print.

What You'll Own
  • Your work will be tracked and followed by the likes of Elon Musk, Mark Zuckerberg, Alexandr Wang, Demis Hassabis, Andrew Ng, Amjad Massad, and more
  • Design genuinely hard and useful evaluations that measure frontier model performance on real-world tasks, and that become industry-leading gold-standards
  • Investigate model failures and identify what they reveal about emerging capabilities
  • Publish research, technical reports, and analyses that shape how frontier models are evaluated
What We're Looking For
  • Strong STEM background. You studied Computer Science, Data Science, Statistics, Math, Engineering, Physics, or a related field.
  • Deep curiosity about frontier AI models. You're excited by understanding model behavior, discovering areas of failure, and building better ways to evaluate models.
  • Genuine thirst and intellectual to be on the frontier of AI development.
  • Fearlessness to roll up your sleeves, get your hands dirty, and do real work.
Details
  • Location: San Francisco, Levi's Plaza. We sponsor visas and handle relocation.
  • Work Schedule: Sunday-Friday. Saturdays are yours!
  • Compensation: Competitive salary + meaningful equity. You'd be joining at the stage when ownership matters most.

Similar Jobs

More Jobs at Arcada Labs

More Consumer Technology Jobs

Find similar Member of Technical Staff, Frontier Evals jobs: