Research Scientist - Vision Foundation Models

Epsilon Health

$150K — $180K *
Healthcare
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 6+ years of experience in computer vision or machine learning
  • Expertise in vision encoder model training (e.g. ViT, ConvNeXt)
  • Experience with volumetric or spatiotemporal data
  • Proven ability to implement and adapt complex models
  • Proficient in PyTorch or JAX for distributed training
  • Hands-on experience with radiology-related medical imaging applications

Responsibilities

  • Design, train, and scale vision foundation models for radiology applications
  • Extend 2D pretraining to volumetric data, addressing complex data challenges
  • Evaluate model performance using academic benchmarks and live data
  • Contribute to all stages of model development, from dataset curation to production deployment
  • Stay updated with cutting-edge research in computer vision and medical imaging
  • Drive research excellence through publications and technical blogs

Benefits

  • Opportunity to work with one of the largest medical imaging datasets
  • Engagement in cutting-edge AI-assisted diagnosis research
  • Collaborative environment with a focus on innovation and excellence
  • Potential for significant impact on clinical deployment of AI models
  • Access to professional development and conference opportunities
Full Job Description
Role Overview

We're seeking a Research Scientist with deep expertise in vision foundation models to join our ML Research team. You'll be at the forefront of developing and deploying state-of-the-art vision models for medical imaging applications. This role focuses on pretraining and scaling vision encoders for radiology diagnosis across X-ray, CT, and MRI, with a growing emphasis on 3D volumetric modeling. You'll work with one of the largest and most diverse medical imaging datasets in the industry, pushing the boundaries of what's possible in AI-assisted diagnosis while maintaining the rigor required for clinical deployment.

Key Responsibilities
  • Design, train, and scale vision foundation models for radiology applications across X-ray, CT, and MRI modalities, implementing self-supervised, contrastive, masked image modeling, and joint-embedding predictive (JEPA) frameworks.
  • Extend 2D pretraining recipes to volumetric CT and MR data, addressing long sequence lengths, anisotropic spacing, and multi-sequence studies.
  • Evaluate model performance rigorously across academic benchmarks, internal offline datasets, and live production data.
  • Contribute hands-on to all stages of model development including dataset curation, architecture design, distributed training, and production deployment.
  • Stay current with cutting-edge research in computer vision and medical imaging AI.
  • Drive research and technical excellence through conference publications and technical blog posts, establishing best practices for training robust medical imaging models at scale.
Qualifications
  • 6+ years of academia/industry experience in computer vision/machine learning
  • Deep expertise in training vision encoder models at scale (e.g. ViT, ConvNeXt). Strong foundation in self-supervised pretraining, including contrastive, masked image modeling, self-distillation, and JEPA-style objectives.
  • Experience training on volumetric or spatiotemporal data (video, 3D medical imaging)
  • Track record of implementing complex models from research papers and adapting them to new domains
  • Proficiency in PyTorch or JAX, with experience training models on multi-GPU/distributed systems
  • Hands-on experience with medical imaging applications, particularly radiology (X-ray, CT, MRI)
  • Strong software engineering skills and ability to write production-quality code
Preferred Qualifications
  • Publications at top-tier conferences (CVPR, ICCV/ECCV, NeurIPS, ICLR, MICCAI)
  • Experience with 3D medical image processing and retrieval tasks
  • Familiarity with CT and MR acquisition (windowing, multi-sequence protocols, voxel spacing)
  • Experience with long-context training techniques (sequence parallelism, efficient attention)
  • Knowledge of vision-language models and multimodal learning
  • Experience with model interpretability and explainability methods
  • Understanding of clinical evaluation metrics, clinical workflows, and healthcare data (DICOM, HL7, etc.)

Similar Jobs

More Jobs at Epsilon Health

More Healthcare Jobs

Find similar Research Scientist - Vision Foundation Models jobs: