Bachelor's degree in Computer Science, Computer Engineering, or related field (Master's preferred)
4+ years of experience in backend development with Python, Golang, or similar languages
Experience in building multi-modal data pipelines with Ray or Apache Airflow
Proficiency in distributed computing using Spark or Hadoop
Familiarity with relational databases and RESTful APIs
Basic understanding of cloud infrastructure (AWS, GCP)
Exposure to AI/ML concepts preferred
Responsibilities
Develop and maintain scalable backend systems for high-performance AI applications in healthcare
Collaborate with data scientists and ML engineers to design efficient data pipelines for healthcare datasets
Build and manage APIs and microservices for data retrieval and AI model interaction
Monitor and optimize backend systems for performance and reliability
Work closely with product managers to translate healthcare requirements into technical solutions
Develop infrastructure supporting data ingestion, feature extraction, and tagging
Implement tag management and metadata systems for dataset organization
Benefits
Environment for innovation in healthcare AI
Collaboration with physicians and cross-functional teams
Opportunity to work with advanced AI technologies
Focus on building impactful solutions for patient outcomes
In-person team dynamic in Menlo Park, fostering strong collaboration
Full Job Description
Senior Research Engineer
About the Role
Hippocratic AI's AI agents are already working inside real hospitals and health systems - this role owns the backend infrastructure that keeps them fast, reliable, and ready to scale as that footprint grows. You'll architect systems built for tomorrow's volume, not just today's, working closely with data scientists, ML engineers, and product managers to turn healthcare requirements into production-grade infrastructure.
What You'll Do
Build Scalable, Reliable AI Infrastructure
Architect backend systems that sustain 99.9%+ uptime for high-volume healthcare data and LLM processing as usage grows exponentially
Implement monitoring that surfaces issues before they reach production AI agents
Own performance and reliability improvements across backend systems in collaboration with data scientists and ML engineers
Engineer Data Pipelines for Healthcare AI
Design data pipelines that ingest, process, and prepare large-scale, multi-modal (speech, vision, text) healthcare datasets for training and inference with minimal latency
Build tag management and metadata systems that make large datasets organized and retrievable
Develop and optimize infrastructure supporting data ingestion, feature extraction, and tagging workflows
Ship APIs and ML Enablement
Develop APIs and microservices that reduce processing time and improve responsiveness for AI model interaction and data retrieval
Build infrastructure that makes ML workflows reproducible end-to-end, from data preparation through model deployment
Accelerate time-to-production for new AI capabilities
Location and Travel
We believe the best ideas happen together. To support fast collaboration and a strong team culture, this role is expected to be in our Menlo Park office five days a week.
What You Bring Must-Have
Bachelor's degree in Computer Science, Computer Engineering, or a related field (Master's preferred)
4+ years of backend development experience using Python, Golang, or similar languages
Experience building and maintaining multi-modal (speech, vision, text) data pipelines using Ray, Apache Airflow, or similar for distributed processing and model experimentation, including distributed computing frameworks like Spark or Hadoop
Familiarity with relational database systems and RESTful APIs
Basic understanding of cloud infrastructure (AWS, GCP, or similar)
Nice-to-Have
Exposure to AI/ML concepts or experience working with LLMs
Experience working in teams that handle sensitive or regulated data
Familiarity with gRPC, GraphQL, or similar
Experience with real-time audio
Exposure to DevOps concepts - CI/CD, deployment, Terraform, build systems
Experience with data science
What Success Looks Like
Within 6-12 months, the backend systems you've built or hardened are sustaining 99.9%+ uptime under exponential growth in healthcare data and LLM traffic. New AI capabilities move from data preparation to production faster because the ML workflows and pipelines you designed are reproducible end-to-end. Your monitoring catches issues before they ever reach a live AI agent in a hospital setting.