Senior Machine Learning Engineer, AI Evaluation

Society for Human Resource Management (SHRM)

$100K — $130K *
Business Services
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science, Data Science, Machine Learning, Engineering, or a related field; Master's degree preferred.
  • Minimum 7 years of experience in ML/LLM engineering or applied AI, particularly in production systems.
  • Hands-on experience with multi-model LLM setups, including APIs and evaluation frameworks.
  • Proficient in AI model evaluation and benchmarking, including nuances of measurement techniques.
  • Experience with cloud-based infrastructure and data management, preferably on Google Cloud Platform.

Responsibilities

  • Design and maintain engineering infrastructure for structured AI evaluations and experiments.
  • Develop a unified orchestration layer for consistent evaluations across different AI model providers.
  • Implement robust scoring frameworks and methodologies for evaluating AI models' performance.
  • Build reproducibility systems for experimentation with comprehensive logging and tracking.
  • Collaborate with HR experts to translate evaluation standards into technical specifications.
  • Develop structured data repositories for storing evaluation results and monitor performance over time.

Benefits

  • Professional growth and development opportunities.
  • Health, dental, and vision insurance.
  • Health savings and flexible spending accounts.
  • Retirement plan options.
  • Generous open leave policy and annual discretionary bonuses.
Full Job Description
Summary

The Senior Machine Learning Engineer, AI Evaluation builds and operates the measurement and engineering infrastructure supporting the organization's Applied AI Research (AAIR) function, a continuous experimental environment designed to evaluate how artificial intelligence models perform real-world HR and workplace-related tasks against established professional standards.

This role is responsible for designing and maintaining the engineering infrastructure used to conduct rigorous, reproducible AI model evaluations and benchmarks. The Senior Machine Learning Engineer develops the systems that run multiple AI models against structured, domain-specific evaluations; builds scoring and evaluation frameworks; maintains reproducibility across model versions; and creates the data infrastructure necessary to analyze and track model performance over time.

Working closely with HR subject matter experts and Applied AI Research colleagues, this position translates expert-defined standards and evaluation criteria into technically rigorous, measurable specifications. Subject matter experts establish the domain-specific ground truth and standards for correctness, while the Senior Machine Learning Engineer owns the technical systems and methodologies used to measure model performance against those standards.

The position serves as a shared technical engineering resource across multiple Applied AI Research workstreams and helps ensure that published findings, benchmarks, and research conclusions are supported by reliable, auditable, and defensible measurement practices.

This is an AI evaluation and engineering infrastructure role rather than a model-training or frontier AI research position The role is not responsible for developing novel model architectures or training foundation models.

Responsibilities

AI Evaluation Engineering & Infrastructure
  • Design, build, and maintain scalable engineering infrastructure for conducting structured evaluations and experiments across multiple AI and large language model (LLM) families.
  • Develop and maintain a unified, provider-agnostic orchestration layer that enables consistent evaluation across multiple frontier model providers and architectures.
  • Design and implement rigorous AI evaluation and scoring frameworks, including rubric-based scoring, model-as-judge methodologies with appropriate safeguards, partial-credit methodologies, and approaches for managing ambiguity.
  • Build systems and processes that support reproducible experimentation, including model-version pinning, comprehensive run logging, experiment tracking, and drift detection.
  • Maintain portable evaluation architecture across AI model providers to enable consistent and defensible cross-model comparisons as models and technologies evolve.
  • Establish and maintain technical standards and engineering practices that support reliable, repeatable, and auditable AI evaluation.

Evaluation Methodology & Applied AI Research
  • Partner closely with HR subject matter experts to translate professional standards, research criteria, and judgment-based rubrics into measurable and technically executable evaluation specifications.
  • Identify and surface ambiguity, inconsistencies, or measurement limitations within proposed evaluation criteria and collaborate with subject matter experts to strengthen evaluation design.
  • Apply knowledge of AI evaluation methodologies, benchmarking techniques, inter-rater reliability, and known limitations of automated and model-as-judge evaluation approaches.
  • Support the design of measurement methodologies when definitive ground truth is unavailable or requires expert interpretation.
  • Ensure evaluation methodologies align with research-defined validation standards and produce findings that are reproducible, transparent, and defensible.
  • Contribute technical expertise to the design and continuous improvement of AI research experiments, benchmarks, and evaluation methodologies.

Data, Monitoring & Reproducibility
  • Design and maintain structured repositories for experiment results, prompt libraries, scoring rubrics, model metadata, and longitudinal evaluation data using cloud-based data infrastructure.
  • Build and maintain data structures in BigQuery or comparable platforms that enable research results to be queried, analyzed, reproduced, and audited.
  • Develop monitoring, reporting, and visualization capabilities using Looker, Looker Studio, or comparable tools to provide visibility into experiment status, model performance, and performance drift.
  • Establish processes for tracking changes in model behavior across model versions and over time.
  • Maintain complete technical documentation and metadata necessary to reproduce research findings and evaluation results.
  • Ensure appropriate quality controls are incorporated throughout data collection, evaluation, scoring, storage, and reporting processes.

Cross-Functional Technical Partnership
  • Serve as a shared engineering resource across multiple Applied AI Research teams and research workstreams.
  • Collaborate with research leaders, HR subject matter experts, data professionals, engineers, and other internal stakeholders to translate research requirements into scalable technical solutions.
  • Support collaboration with university, affiliate, research, and other external partners when appropriate and within established organizational access controls and data-handling requirements.
  • Communicate technical methodologies, limitations, risks, and findings clearly to both technical and non-technical audiences.
  • Evaluate emerging AI models, tools, technologies, and evaluation methodologies and recommend appropriate applications within the research environment.
  • Contribute to continuous improvement of the Applied AI Research technical environment, engineering practices, and evaluation capabilities.

Data Governance, Security & Responsible AI
  • Ensure AI evaluation systems and workflows comply with organizational requirements for data security, privacy, ownership, access, and responsible AI use.
  • Maintain appropriate controls to protect proprietary, member, research, and other sensitive data from unauthorized access or use.
  • Ensure organizational data is not used to train external shared models except where expressly authorized and appropriately governed.
  • Implement and maintain appropriate access controls and data-handling requirements when working with external research or partner organizations.
  • Partner with appropriate internal stakeholders to ensure evaluation infrastructure aligns with organizational technology, security, privacy, and governance standards.

Education & Experience Requirements

Education
  • Bachelor's degree in Computer Science, Data Science, Machine Learning, Engineering, or a related quantitative or technical field, or relevant equivalent experience in lieu of degree.
  • Master's degree in Computer Science, Data Science, Machine Learning, Artificial Intelligence, or a related field preferred.
  • A Ph.D. is not required; demonstrated expertise in AI/ML evaluation engineering, research infrastructure, and production-grade systems is valued.

Experience
  • Seven (7) or more years of progressively responsible experience in ML/LLM engineering, applied AI, applied data science, or research infrastructure, including experience developing, implementing, and supporting production-grade systems.
  • Demonstrated hands-on experience developing multi-model LLM applications and infrastructure, including provider-agnostic model access, APIs, prompt engineering, and evaluation frameworks.
  • Demonstrated experience with AI model evaluation and benchmarking, including rubric-based scoring, inter-rater reliability, model-as-judge methodologies and their limitations, and approaches for evaluating performance when definitive ground truth may not be readily available.
  • Experience designing and maintaining reproducible technical systems incorporating version control, model-version pinning, comprehensive logging, experiment tracking, and/or drift detection.
  • Experience with cloud-based data and AI infrastructure on a major cloud platform; Google Cloud Platform experience, including BigQuery, Vertex AI, IAM, and audit logging, preferred.
  • Experience developing and supporting data pipelines, structured experiment repositories, dashboards, or monitoring solutions.
  • Experience working with HR, workforce, survey, behavioral, or other professional-domain data preferred.
  • Experience supporting academic, applied research, benchmarking, or peer-reviewed research workflows preferred.

Certifications

Knowledge, Skills & Abilities
  • Advanced proficiency in Python and strong software-engineering fundamentals, including the ability to develop reliable, maintainable, production-quality code.
  • Strong knowledge of machine learning, large language models, generative AI systems, and contemporary AI application architectures.
  • Demonstrated knowledge of AI evaluation and benchmarking methodologies, including rubric-based evaluation, automated scoring, model-as-judge approaches, inter-rater reliability, and measurement design.
  • Strong understanding of the limitations and failure modes of generative AI systems and the ability to design evaluation approaches that appropriately account for those limitations.
  • Demonstrated commitment to reproducibility, including disciplined use of versioning, documentation, logging, experiment tracking, and drift detection.
  • Ability to translate complex, judgment-based requirements from subject matter experts into technically rigorous and measurable evaluation specifications without oversimplifying the underlying domain expertise.
  • Strong analytical and problem-solving skills with the ability to identify technical, methodological, and data-quality issues and develop appropriate solutions.
  • Working knowledge of cloud-based AI and data environments, APIs, data warehouses, access controls, and related technical infrastructure.
  • Ability to effectively communicate complex technical concepts, methodologies, limitations, and findings to technical and non-technical audiences.
  • Strong collaboration and consultation skills, with the ability to work effectively with researchers, engineers, data professionals, subject matter experts, and external partners.
  • Ability to balance technical rigor, research requirements, scalability, and practical implementation considerations.
  • Strong understanding of responsible AI principles, data governance, privacy, security, and appropriate handling of sensitive or proprietary information.
  • Ability to evaluate emerging AI models, technologies, and evaluation methodologies and determine their appropriate application within the organization's research environment.
  • Ability to effectively leverage AI tools and technologies to streamline workflows, enhance productivity, and improve overall work quality.


Physical Requirements

This position operates in a typical office environment (which includes a home office setting) and requires the ability to perform essential job functions with or without reasonable accommodation. Physical requirements may include:
  • Prolonged periods of sitting at a desk and working on a computer.
  • Frequent use of hands and fingers for typing, handling documents, and using office equipment.
  • Occasional standing, walking, bending, and reaching.
  • Ability to lift and carry up to 30 pounds as needed.
  • Clear verbal and written communication skills for effective interaction with colleagues and stakeholders.

Work Environment

Hybrid Schedule (3 Days In-Office/2 Days Remote)

This position follows a hybrid work schedule, with Tuesday through Thursday in office and Monday and Friday remote. Employees must be available during standard business hours, with core hours beginning between 8:00-9:00 a.m. and concluding between 5:00-6:00 p.m. local time.

Travel: Occasional 0 - 10%

#LI

The hiring range for this position is $100,000 to $130,000 per year. This range is an estimate, and the actual salary may vary based on the candidate's experience, skills, and qualifications. SHRM offers a competitive and comprehensive total rewards package. The benefits for this position include professional growth and development, health, dental, vision, well-being, health savings, flexible spending, retirement, open leave, and annual discretionary bonus and incentives.

Similar Jobs

More Jobs at Society for Human Resource Management (SHRM)

More Business Services Jobs

Find similar Senior Machine Learning Engineer, AI Evaluation jobs: