Role: Software Engineer, Machine Learning Operations
Pod: Machine Learning
Location: Hayes Valley, San Francisco, CA
Basic Job DetailsJob Type: Full Time
Work Model: Hybrid
Remote Days: Monday and Friday
Office Days: Tuesday, Wednesday, and Thursday
Job DescriptionAs a Software Engineer on Baton's Machine Learning Pod, you will build and maintain the production infrastructure that supports the full machine-learning lifecycle. You will work across production software engineering, distributed systems, MLOps, and model development to help the team bring new models online and operate them reliably at scale.
Baton's primary ML infrastructure is established, and the team is now building the next layer of MLOps capabilities on top of that foundation. You will help automate model monitoring, retraining, redeployment, experimentation, and drift detection as the number of production models continues to grow.
This is a hands-on individual contributor role for an engineer who can work across both infrastructure and modeling. You will build on the patterns and templates the team has already established, improve integration between the ML platform and Baton's core transportation management platform, and make it easier for engineers to develop, ship, and maintain models end to end.
Responsibilities- Build and Expand MLOps Infrastructure:
- Build automated capabilities for model monitoring, retraining, redeployment, champion/challenger testing, A/B testing, and drift detection.
- Improve experiment tracking and model lifecycle management as the number of production models increases.
- Develop and Productionize Machine-Learning Models:
- Bring new machine-learning models into production, including developing select models from initial concept through deployment.
- Support models across development, deployment, monitoring, maintenance, and iteration.
- Build scalable batch-prediction capabilities alongside real-time machine-learning workflows.
- Create Self-Serving ML Infrastructure:
- Build on existing infrastructure patterns and templates to create reliable and reusable ML workflows.
- Make it easier for engineers to ship and maintain models end to end with less manual intervention.
- Improve development velocity while maintaining production reliability and operational quality.
- Strengthen Distributed ML Systems:
- Design and maintain distributed systems that support data-intensive and machine-learning workloads.
- Improve the scalability, performance, and reliability of production ML infrastructure.
- Contribute to batch processing, caching, data movement, and cloud-native infrastructure.
- Connect ML Systems with Baton's Core Platform:
- Strengthen the integration between the ML platform and Baton's core transportation management platform.
- Replace manual integration workflows with scalable and maintainable infrastructure.
- Enable machine-learning capabilities to support transportation workflows and operational decision-making.
- Collaborate Across the ML Lifecycle:
- Partner with engineers and cross-functional stakeholders to identify opportunities for automation and model productionization.
- Contribute across software engineering, ML development, infrastructure, and production operations based on the needs of the team.
Required QualificationsProduction Python Expertise- Advanced proficiency coding in production-grade Python at an L4 or L5 level
- Experience working in an environment where production code directly impacts operations
- Ability to build and maintain reliable software across modeling, infrastructure, and automation workflows
Distributed Systems Expertise- Strong background in distributed computing, scalable ML infrastructure, and high-performance engineering
- Experience building or maintaining systems that support data-intensive and ML workloads
- Familiarity with big-data systems, batch processing, caching, and cloud infrastructure
Machine Learning / MLOps- Experience implementing, deploying, and productionizing machine-learning algorithms
- Hands-on experience with data engineering, distributed training, model monitoring, and experiment tracking
- Experience with model retraining, redeployment, serving, and lifecycle management
- Strong SQL knowledge and caching experience
- Experience with model lifecycle platforms such as SageMaker is a plus and should be confirmed with Fabian as a must-have versus preferred qualification
Preferred Qualifications- Experience implementing, deploying, monitoring, and maintaining machine-learning models in production.
- Experience with Kubernetes and cloud infrastructure, preferably AWS.
- Familiarity with ML and data technologies such as Kubeflow, Iceberg, Feast, or SageMaker.
- Experience with batch prediction, model serving, distributed training, experiment tracking, caching, or feature stores.
- Experience building scalable, self-serving infrastructure for machine-learning teams.
- Experience integrating ML platforms with broader production or operational systems.
- Previous experience in a technically rigorous environment such as a large-scale technology company, infrastructure organization, or high-growth engineering team.
- Experience in logistics, transportation, freight, or supply chain is a plus but not required.
The Perks- Competitive Base Salary + Cash Bonus Structure
- Annual Company Bonus + Long Term Incentive Plan
- 401(k) with Matching
- Hybrid Work Schedule
- Hyper-Stable, Publicly Traded Enterprise
- Medical, Dental, and Vision Health Coverage
- Employee Stock Purchase Program with a 15% Discount to Market Value
- Collaborative, Fun, and Tech-Forward Office in Hayes Valley, San Francisco
Compensation Range: The annual base salary range for this position is $162,000 - $216,000*
Compensation will vary based on factors including skill level, transferable knowledge, and experience.
Note that the above is not the representation of total compensation, which includes our LTI Package as well.
In addition to base salary, Baton's full-time employees are eligible for an annual company performance bonuses.