About the roleAs a Research Engineer focused on Model Evaluation & MLOps, you will build the tools and infrastructure needed to evaluate, deploy, and operate multimodal foundation models reliably. You will rapidly enable Sciforium's models and the latest open-weight models on GPUs, automate quality and performance benchmarking, and improve the MLOps workflows that connect research experiments to reliable releases.
Key ResponsibilitiesModel Enablement & Automated Evaluation- Rapidly integrate new internal and open-weight language and multimodal models into our GPU evaluation and inference environments.
- Build automated benchmarks for model quality and systems performance, including latency, throughput, and memory usage.
- Create standardized, reproducible comparisons across Sciforium models, external baselines, and runtime configurations.
MLOps & Model Lifecycle- Build and maintain experiment tracking, model registry, and versioning for models, datasets, and evaluation configurations.
- Automate the path from research checkpoints to validated deployments through CI/CD and reproducible workflows.
- Monitor model quality and systems performance, and diagnose failures or regressions across model and deployment pipelines.
Research & Systems Collaboration- Build reusable tools that help researchers launch evaluations, compare experiments, and reproduce results.
- Profile end-to-end model workloads and collaborate with distributed systems, inference, and GPU kernel engineers on deeper performance issues.
Must-HavesCandidates may be stronger in some areas than others. We are looking for strong software engineering foundations, hands-on ML systems experience, and depth in at least one of model evaluation, MLOps, or model deployment.
- Experience: 2+ years of professional ML or software engineering experience, including work on production ML systems, ML platforms, or MLOps infrastructure.
- Software Engineering: Strong Python and software engineering skills, with experience building reliable production systems.
- Machine Learning Expertise: Hands-on experience with PyTorch, TensorFlow, or JAX and a good understanding of modern language or multimodal model architectures.
- Evaluation & MLOps: Experience with model evaluation or benchmarking and core model lifecycle workflows such as experiment tracking, versioning, deployment, or monitoring.
- GPU Systems: Experience running, benchmarking, and debugging models with one or more GPU inference runtimes, such as vLLM, SGLang, TensorRT-LLM, or equivalent, in containerized cloud or on-premises environments.
- Communication: Ability to document systems clearly and collaborate across research, infrastructure, and product engineering teams.
- Education: MS or PhD in Computer Science, Computer Engineering, Machine Learning, or a related technical field, or equivalent practical experience.
Nice-to-Have- Familiarity with Hugging Face Transformers or similar model libraries.
- Experience enabling models on AMD GPUs and ROCm.
- Contributions to open-source evaluation, model, or ML infrastructure projects.
Benefits include- Medical, dental, and vision insurance
- 401k plan
- Daily lunch, snacks, and beverages
- Flexible time off
- Competitive salary and equity