Member of Technical Staff - Research

Crucibl

$150K — $180K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Published research at a top venue or research role at a frontier lab required
  • Deep understanding of LLM reasoning and evaluation crucial
  • Strong fundamentals in ML, statistics, or NLP necessary
  • Ability to quickly transition from research questions to testable hypotheses essential

Responsibilities

  • Research ambiguous business decisions made by frontier models and identify failures
  • Design innovative reasoning and evaluation methods that exceed standard benchmarks
  • Translate open research problems into testable approaches for application
  • Define evaluation frameworks for measuring 'good judgment' in models
  • Develop experiments to uncover real model failure modes
  • Provide concrete recommendations for product and applied teams
  • Collaborate with founders to shape technical vision and roadmap

Benefits

  • Opportunity to influence product, culture, and company direction in a seed-stage startup
  • Clear impact of work; see tangible results from contributions
  • Join a small, exceptional team that emphasizes hiring the best talent
  • Comprehensive salary and equity package included in every offer
  • Flexible hybrid work schedule with a focus on results, not just presence
  • Complete medical, dental, and vision coverage provided
  • Daily lunch and snacks offered to keep you fueled
Full Job Description
What You'll Do

Push the Frontier of Judgment
  • Research how frontier models reason through ambiguous, high-stakes business decisions - and where they fail
  • Design novel methods for reasoning, evaluation, and calibration that go beyond standard benchmarks
  • Translate open problems in reasoning, uncertainty, and multi-step decision-making into approaches we can test and ship


Build Evaluation That Matters
  • Define what "good judgment" looks like for a model, and build the evaluation frameworks to measure it
  • Design experiments that reveal real failure modes, not just leaderboard scores
  • Turn findings into concrete recommendations for the product and applied teams


Partner with Founders
  • Shape technical vision and roadmap alongside the founding team
  • Bring outside research thinking into a company solving a problem few labs are focused on


Set the Bar
  • Define what rigorous, applied research looks like in an AI-first organization
  • Raise the bar for the team as it grows


Who You Are Must-Haves
  • At least one undeniable signal of excellence - published research at a top venue, research role at a frontier lab, or a track record of novel technical contributions
  • Deep understanding of how LLMs reason, fail, and can be evaluated - this isn't a theoretical interest, it's core to the job
  • Strong fundamentals in ML, statistics, or NLP - we're too small for hand-holding on the technical side
  • Comfortable moving from open-ended research question to a testable hypothesis quickly


Strong Signal
  • Experience designing evaluation frameworks or benchmarks for reasoning, decision-making, or agentic systems
  • Published or shipped work on uncertainty, calibration, multi-step reasoning, or LLM evaluation
  • Has operated in a high-growth, high-ambiguity environment
  • Thinks like an owner - big picture, not just your part


We're Not Looking For
  • Researchers who need a clean, self-contained problem before starting - ambiguity is the job
  • Pure benchmark-chasers - we care about judgment that holds up with real clients, not just leaderboard gains
  • Researchers without product instincts - your work needs to change what we build and ship, not sit in a paper -- but we are open to publishing work!

Who You'll Work With

Our founders bring deep technology experience (Google-scale systems serving hundreds of millions of users) and domain expertise in high-stakes business decisions from top consulting and private equity firms. You'll be the connective tissue between these worlds, and between the company we are today and the company we're becoming.

Benefits & Perks
  • Zero to Outcome: Seed stage means you shape the product, the culture, and the trajectory - not inherit them.
  • See Your Work Matter: No abstractions, no layers - you'll see exactly what you built and what it changed.
  • The People Around You: A small, elite team that will raise your game. We hire for exceptional, not just experienced.
  • Meaningful Equity: Every offer includes a comprehensive salary and equity package.
  • Hybrid Schedule: 3 days in-office with real flexibility around the rest. We care about output, not optics.
  • Full Health Coverage: Medical, dental, and vision.
  • Daily Lunch & Snacks: Fueled and focused, on us.

Work Authorization

Crucibl welcomes applications from candidates requiring visa sponsorship. Sponsorship eligibility is determined during the interview process.

Similar Jobs

More Jobs at Crucibl

  • Member of Technical Staff
    $150K — $180K *
    San Francisco, CA 94112 (San Francisco County)
    Enterprise Technology
    In-Person
  • AI Strategist
    $130K — $160K *
    San Francisco, CA 94112 (San Francisco County)
    Business Services
    In-Person
  • Founding Product Manager
    $130K — $160K *
    San Francisco, CA 94112 (San Francisco County)
    Enterprise Technology
    In-Person

More Information Technology Jobs

Find similar Member of Technical Staff - Research jobs: