AI Evals Engineer - Evaluation Datasets & Ground Truth

Prophetic Technologies Inc

$110K — $130K *
US-AnywhereRemote in United States
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience with ML or LLM systems in production, focusing on output quality.
  • Proven experience in building evaluation datasets, with a deep understanding of how to define correctness.
  • Fluency in ML validation principles, including sampling techniques and precision/recall trade-offs.
  • Strong skills in Python and SQL for data manipulation and analysis.
  • Hands-on familiarity with LLMs, including prompting and understanding their failure modes.
  • Ability to independently translate system designs into effective label schemas.

Responsibilities

  • Define the metrics and measurements needed for evaluable modules in collaboration with product and engineering.
  • Establish and write clear definitions for correctness including rubrics and label schemas.
  • Prioritize validation tasks, identifying which will most significantly impact product iteration.
  • Source reliable data, utilizing strategies like stratified sampling and human labeling.
  • Ensure data quality by managing version control and guarding against contamination over time.
  • Collaborate closely with engineering to calibrate grading systems and ensure your datasets are effectively utilized.

Benefits

  • 100% coverage for medical, dental & vision insurance for employees, with 30% for dependents.
  • Unlimited paid time off
  • A hybrid/remote work stipend to support flexible work environments.
  • Snacks and drinks available in the office, plus a friendly office Doberman for morale.
  • Budget allocated for intra-office travel and team meetups.
Full Job Description
Prophetic Software is not able to sponsor employment visas now or in the future. Candidates must be authorized to work in the United States without current or future sponsorship to be considered for this role.

Why this role exists

We iterate on pipeline components constantly: prompts, models, classifier thresholds, retrieval logic, harness design. We are hiring a full-time engineer to produce, maintain, and defend the evaluation datasets that let every meaningful component in our stack be measured. You will own ground truth at Prophetic.

This is not an evals-infrastructure role and it is not a labeling-operations role, although you'll touch both. Your deliverable is trusted data: for a given module, a versioned set of inputs and expected outputs, plus a written definition of what "correct" means and how confident we should be in the labels.

What you'll do

Decide what needs to be measured, and how
  • Read our pipelines and system architecture, sit with product and engineering, and decompose each system into evaluable modules with explicit input 12 expected-output contracts.
  • For each module, define what "correct" means in writing: rubrics, label schemas, edge-case policies, and the tolerances that matter to the business.
  • Prioritize. We have more modules than you can cover in year one; you'll decide where a validation set unblocks the most iteration, with minimal direction from principal engineers.

Source the data by whatever means is cheapest and most trustworthy for that module
  • Production sampling: pull stratified, de-identified samples from real traffic so eval sets reflect what the system actually sees, including the long tail.
  • Human labeling: scope and run labeling programs, write annotation guidelines, build calibration sets, measure inter-annotator agreement, and manage vendors (including offshore labeling teams) or internal subject-matter experts. You own label quality, not just label throughput.
  • Synthetic / oracle-generated ground truth: where a task is tractable for a frontier model given enough compute (long context, multi-pass, tool use, self-consistency) but too expensive to run that way in production, design the oracle harness that produces labels for the cheap production path to be measured against. Then verify the oracle: calibrate its output against a human-labeled sample before anyone trusts it.
  • Programmatic and adversarial construction: heuristic labels, templated edge cases, backtests built from past incidents ("what test would have caught this?").

Make the data trustworthy over time
  • Version every dataset; track lineage, splits, and which model/prompt versions have seen which examples.
  • Guard against contamination and leakage (eval examples drifting into few-shot prompts, the oracle model also being the production model, etc.).
  • Slice by customer segment, input type, and difficulty so a headline number can't hide a regression.
  • Refresh sets as the product and traffic change; retire stale examples.

Close the loop with engineering
  • Calibrate automated graders (LLM-as-judge, similarity metrics, exact-match) against your human gold sets, and be the person who says when an automated judge is good enough to gate on.
  • Report metrics correctly: precision/recall/F1, confusion matrices, calibration, confidence intervals, sample sizes needed to detect a given effect.
  • Partner with engineers who own the eval harness and CI so your datasets are actually run, and with ML engineers on classifier features when the data tells you the features are the problem.
What we're looking for

Must have
  • Experience building or evaluating ML or LLM-powered systems in production, in a role where output quality was your problem.
  • You have built evaluation or validation datasets before and can talk about one in detail: how you defined correctness, how you sourced labels, what went wrong, how you knew the labels were good.
  • Working fluency in ML validation fundamentals: train/validation/test discipline, stratified sampling, precision/recall trade-offs, class imbalance, calibration, inter-rater agreement, basic significance testing and power.
  • Understand feature engineering well enough to reason about why a classifier fails and what data would expose it.
  • Strong Python and SQL; comfortable pulling and reshaping data yourself.
  • Hands-on with LLM-based systems: prompting, structured outputs, agent/tool-use harnesses, and the specific ways they fail (non-determinism, prompt sensitivity, evaluator bias).
  • Judgment about when LLM-as-judge is reliable and when it is not, and how to prove either.
  • You can read a system design, understand the business logic it encodes, and translate that into a label schema without waiting to be told.

Nice to have
  • Have run a human-labeling program end to end, including vendor selection, guideline authoring, QA sampling, and cost/quality trade-offs.
  • Experience with eval tooling.
  • Experience with data labeling platforms.
  • Have used a frontier model as a distillation/oracle source and can articulate where that assumption breaks.

How you think
  • Verify the verifier. A label is a claim, not a fact, until something independent agrees with it.
  • Start simple: binary before graded, one judge before five, a hundred well-understood examples before ten thousand noisy ones.
  • You measure business outcomes, not model vibes, and you're comfortable telling a senior engineer their favorite change didn't move the number.
  • You'd rather own an unglamorous problem completely than a glamorous one partially.
Team & reporting

You'll report to the Chief AI Officer and work across all product/pipeline teams. Over time it is expected that you'll lead a small team of eval engineers. You'll have a generous budget for labeling vendors and oracle compute.

Benefits (US-based Full-time):
  • 100% medical, dental & vision insurance coverage for you; 30% coverage for dependents
  • Competitive salary and meaningful early-stage equity
  • Unlimited PTO
  • Hybrid/Remote stipend
  • In-office perks: snacks, drinks, coffee, ping-pong table, and more - plus cuddles from Olive, our in-office Doberman!
  • Budget for intra-office travel
  • 2-3 annual team meetups in person


Similar Jobs

More Jobs at Prophetic Technologies Inc

  • Product Manager
    $110K — $130K *
    Remote
    Real Estate & Construction
    Remote
  • Solutions Engineer
    $110K — $130K *
    Portland, OR 97229 (Washington County)
    Enterprise Technology
    In-Person
  • Revenue Operations Manager
    $110K — $130K *
    Portland, OR 97229 (Washington County)
    Business Services
    In-Person
  • Software Engineer
    $90K — $130K *
    Portland, OR 97229 (Washington County)
    Information Technology
    In-Person
  • Account Executive
    $80K — $120K *
    Remote
    Enterprise Technology
    Remote in United States

More Information Technology Jobs

Find similar AI Evals Engineer - Evaluation Datasets & Ground Truth jobs: