Job Summary:
Seeking an MLOps Engineer with 5+ years of experience in MLOps, Machine Learning Engineering, Data Engineering, or Platform Engineering. The ideal candidate will have strong Python development skills, hands-on experience with GCP, Databricks, Airflow, CI/CD, and infrastructure automation, along with enterprise-level experience supporting production machine learning systems. This role focuses on productionizing, scaling, monitoring, and optimizing machine learning solutions while working as an individual contributor in a remote environment.
Key Responsibilities:
• Deploy machine learning models into scalable production environments.
• Build and maintain CI/CD pipelines for machine learning workloads.
• Partner with Data Scientists to operationalize new machine learning models and capabilities.
• Implement model serving and deployment strategies across multiple business units.
• Design and support scalable, multi-tenant machine learning infrastructure.
• Improve platform reliability, performance, observability, and operational consistency.
• Develop reusable deployment frameworks and infrastructure templates.
• Ensure platform consistency across multiple regions and business domains.
• Monitor model health, model performance, data quality, and overall system reliability.
• Establish alerting, logging, and automated remediation processes.
• Optimize infrastructure utilization and cloud costs.
• Troubleshoot production incidents and drive root-cause resolution.
• Build and optimize batch and real-time machine learning pipelines.
• Improve workflow orchestration and data movement processes.
• Support feature engineering pipelines and model retraining frameworks.
• Ensure the reliability and scalability of production data infrastructure.
Required Qualifications:
• 5+ years of experience in MLOps, Machine Learning Engineering, Data Engineering, or Platform Engineering.
• Strong Python development experience.
• Hands-on experience with GCP.
• Hands-on experience with Databricks.
• Hands-on experience with Airflow.
• Experience building and maintaining CI/CD pipelines.
• Experience with infrastructure automation.
• Experience deploying machine learning solutions in cloud environments.
• Experience supporting production machine learning systems at an enterprise level.
• Strong individual contributor experience with the ability to independently own technical responsibilities.
• Strong understanding of machine learning operations, deployment, scalability, reliability, and automation.
Preferred Qualifications:
• Experience supporting search, recommendation, ranking, or advertising platforms.
• Knowledge of model monitoring and observability frameworks.
• Experience with cloud cost optimization and performance tuning.
• Experience supporting multi-tenant platforms.
• Familiarity with modern MLOps frameworks and deployment patterns.