Steampunk

AI Evaluation Scientist

Steampunk$105K — $145K *
Consumer Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's or Master's degree in Computer Science, Statistics, Machine Learning, Cognitive Science, or a related field.
  • 2+ years of experience in evaluating machine learning or generative AI models, especially LLMs.
  • Proficiency in Python and libraries like PyTorch and Hugging Face.
  • Familiarity with evaluation metrics and experimental design for AI systems.
  • Strong analytical skills and ability to present findings clearly to different audiences.

Responsibilities

  • Implement evaluation frameworks to assess AI models on various performance metrics.
  • Build and maintain automated scripts to detect performance drift of AI models.
  • Develop benchmark datasets and scenario-based test cases for mission requirements.
  • Conduct structured error analysis and document recommendations for improvement.
  • Collaborate with engineers and data scientists on model quality and hardening.
  • Design human-in-the-loop evaluation workflows to integrate qualitative insights into reports.
  • Keep updated with emerging tools and research related to AI evaluation.

Benefits

  • Opportunities for professional development and growth in the AI field.
  • Access to cutting-edge evaluation tools and frameworks.
  • Collaborative work environment with cross-functional teams.
  • Support for compliance and risk management efforts in AI deployment.
  • Engagement with innovative AI and data exploitation practices.
Full Job Description
Overview

We are looking for anAI Evaluation Scientistto design and execute evaluation processes that ensure our predictive and generative AI systems areaccurate, reliable, safe, and aligned with mission requirements. This role is essential forestablishingtrust in AI solutions and supporting continuous improvement across the AI lifecycle. The AI Evaluation Scientist will work closely with engineers, data scientists, governance analysts, and product teams to develop evaluation metrics, build test harnesses, analyze model behavior, and support responsible deployment.

Contributions

Responsibilities include:

  • Implement evaluation frameworks for AI models, including accuracy, robustness, relevance, bias, hallucination rate, and safety metrics.
  • Build andmaintainautomated evaluation scripts, tests, and pipelines that assess AI model outputs and detect performance drift over time.
  • Develop benchmark datasets, challenge sets, and scenario-based test cases tailored to mission and user needs.
  • Perform structured error analysis and behavioral audits of LLMs, retrieval-augmented generation (RAG) systems, and predictive models, documentingfindingsand improvement recommendations.
  • Collaborate with AI Developers,LLMOpsEngineers, and Data Scientists to support iterative experimentation, model hardening, and quality improvements.
  • Contribute to the design of human-in-the-loop evaluation workflows, integrating qualitative and quantitative insight into evaluation reports.
  • Assistin mapping evaluation outcomes to responsible AI principles such as fairness, transparency, reliability, and safety.
  • Partner with AI Governance Analysts to ensure evaluation outputs support compliance, documentation, and risk assessments.
  • Stay current with emerging evaluation tools, frameworks, metrics, and research related to LLM assessment and generative AI reliability.
  • Document evaluation processes, criteria, and results for both technical and non-technical audiences.
  • You will contribute to the growth of our AI & Data Exploitation Practice!

Qualifications
  • Ability to hold aposition of public trustwith the U.S. government.
  • Bachelors orMasters degreeinComputer Science, Statistics, Machine Learning, Cognitive Science, Human-Computer Interaction, Data Science, ora relatedfield.
  • 2+ yearsof experience evaluating machine learning models, NLP systems, or generative AI models (LLMs preferred).
  • Familiarity withevaluationmetrics, statistical testing, dataset creation, and experimental design for AI systems.
  • Proficiencyin Python and relevant libraries such asPyTorch, Hugging Face, scikit-learn, LangChain.
  • Proficiency in AI evaluation frameworks such as Ragas.
  • Experience analyzing structured and unstructured data, including text, documents, and embeddings.
  • Understanding ofLLM behavior, prompt evaluation, retrieval pipelines, or RAG architectures.
  • Exposure to responsible AI concepts and governance-aligned evaluation criteria (e.g., fairness, transparency, reliability).
  • Strong analytical skills with the ability to interpret model weaknesses, extract insights, and recommend actionable improvements.
  • Excellent written and verbal communication skills, with the ability to present evaluation findings clearly to technical and non-technical stakeholders.
  • Experience working in agile or iterative development environments is a plus.
  • Familiarity with OWASP LLM Top 10 Risks.
  • Relevant certifications (helpful but not required):
    • NIST AI RMF (AISIC)
    • INFORMS CAP
    • AWS/Azure/Google ML Certifications

About steampunk

Steampunk relies on several factors to determine salary, including but not limited to geographic location, contractual requirements, education, knowledge, skills, competencies, and experience. The projected compensation range for this position is $105,000 to $145,000. The estimate displayed represents a typical annual salary range for this position. Annual salary is just one aspect of Steampunks total compensation package for employees. Learn more about additional Steampunk benefits here.

Identity Statement

As part of the application process, you are expected to be on camera during interviews and assessments. We reserve the right to take your picture to verify your identity and prevent fraud.

Similar Jobs

  • Guidehouse
    Scientist
    $65K — $108K *
    Guidehouse
    Bethesda, MD 20817 (Montgomery County)
  • Guidehouse
    Laboratory Scientist
    $74K — $124K *
    Guidehouse
    Frederick, MD 21702 (Frederick County)
  • Guidehouse
    Biologist
    $113K — $188K *
    Guidehouse
    Bethesda, MD 20817 (Montgomery County)
  • DTRA Biologist
    $130K — $150K *
    Systems Planning And Analysis, Inc.
    Fort Belvoir, VA 22060 (Fairfax County)
  • Defense Contract Audit Agency
    Statistician
    $100K — $120K *
    Defense Contract Audit Agency
    Buffalo, NY 14221 (Erie County)
  • Defense Contract Audit Agency
    Statistician
    $100K — $120K *
    Defense Contract Audit Agency
    Fort Belvoir, VA 22060 (Fairfax County)

More Jobs at Steampunk

More Consumer Technology Jobs

Find similar AI Evaluation Scientist jobs: