Research Engineer - Model Evaluation & MLOps

Sciforium

$135K — $160K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 2+ years of experience in ML or software engineering, focusing on production ML systems or MLOps infrastructure.
  • Strong Python programming and software engineering abilities for building reliable systems.
  • Hands-on experience with frameworks like PyTorch, TensorFlow, or JAX, and understanding of multimodal model architectures.
  • Proven experience in model evaluation, benchmarking, and core workflows like experiment tracking and deployment.
  • Experience with GPU inference runtimes and containerized environments for benchmarking and debugging models.
  • Strong communication skills for effective collaboration across teams and clear documentation.
  • MS or PhD in a related technical field or equivalent practical experience.

Responsibilities

  • Integrate new language and multimodal models into GPU evaluation environments.
  • Develop automated benchmarks for model quality and system performance metrics.
  • Establish reproducible model comparisons across various configurations and baselines.
  • Maintain experiment tracking, model registries, and version control for evaluations.
  • Automate deployment processes through CI/CD and ensure reproducibility of workflows.
  • Monitor model quality and performance, diagnosing issues within deployment pipelines.
  • Create tools for researchers to conduct evaluations and reproduce experimental results.

Benefits

  • Medical, dental, and vision insurance
  • 401k plan
  • Daily lunch, snacks, and beverages
  • Flexible time off
  • Competitive salary and equity
Full Job Description
About the role

As a Research Engineer focused on Model Evaluation & MLOps, you will build the tools and infrastructure needed to evaluate, deploy, and operate multimodal foundation models reliably. You will rapidly enable Sciforium's models and the latest open-weight models on GPUs, automate quality and performance benchmarking, and improve the MLOps workflows that connect research experiments to reliable releases.

Key Responsibilities
Model Enablement & Automated Evaluation
  • Rapidly integrate new internal and open-weight language and multimodal models into our GPU evaluation and inference environments.
  • Build automated benchmarks for model quality and systems performance, including latency, throughput, and memory usage.
  • Create standardized, reproducible comparisons across Sciforium models, external baselines, and runtime configurations.
MLOps & Model Lifecycle
  • Build and maintain experiment tracking, model registry, and versioning for models, datasets, and evaluation configurations.
  • Automate the path from research checkpoints to validated deployments through CI/CD and reproducible workflows.
  • Monitor model quality and systems performance, and diagnose failures or regressions across model and deployment pipelines.
Research & Systems Collaboration
  • Build reusable tools that help researchers launch evaluations, compare experiments, and reproduce results.
  • Profile end-to-end model workloads and collaborate with distributed systems, inference, and GPU kernel engineers on deeper performance issues.
Must-Haves

Candidates may be stronger in some areas than others. We are looking for strong software engineering foundations, hands-on ML systems experience, and depth in at least one of model evaluation, MLOps, or model deployment.
  • Experience: 2+ years of professional ML or software engineering experience, including work on production ML systems, ML platforms, or MLOps infrastructure.
  • Software Engineering: Strong Python and software engineering skills, with experience building reliable production systems.
  • Machine Learning Expertise: Hands-on experience with PyTorch, TensorFlow, or JAX and a good understanding of modern language or multimodal model architectures.
  • Evaluation & MLOps: Experience with model evaluation or benchmarking and core model lifecycle workflows such as experiment tracking, versioning, deployment, or monitoring.
  • GPU Systems: Experience running, benchmarking, and debugging models with one or more GPU inference runtimes, such as vLLM, SGLang, TensorRT-LLM, or equivalent, in containerized cloud or on-premises environments.
  • Communication: Ability to document systems clearly and collaborate across research, infrastructure, and product engineering teams.
  • Education: MS or PhD in Computer Science, Computer Engineering, Machine Learning, or a related technical field, or equivalent practical experience.
Nice-to-Have
  • Familiarity with Hugging Face Transformers or similar model libraries.
  • Experience enabling models on AMD GPUs and ROCm.
  • Contributions to open-source evaluation, model, or ML infrastructure projects.


Benefits include
  • Medical, dental, and vision insurance
  • 401k plan
  • Daily lunch, snacks, and beverages
  • Flexible time off
  • Competitive salary and equity

Similar Jobs

More Jobs at Sciforium

More Information Technology Jobs

Find similar Research Engineer - Model Evaluation & MLOps jobs: