Job Summary
We are seeking an experienced Machine Learning Engineer with strong hands-on expertise in building, training, deploying, monitoring, and maintaining production machine learning models. The role focuses on the end-to-end Machine Learning lifecycle, including data engineering, model development, offline evaluation, production deployment, monitoring, retraining, and A/B testing. The ideal candidate will have strong experience with Python, PySpark, large-scale data platforms, recommendation and ad personalization models, and production ML systems.
Key Responsibilities
Design, build, train, validate, and deploy production Machine Learning models.
Build recommendation and ad personalization models and support offline model evaluation and A/B testing in production.
Perform feature engineering, feature selection, data preprocessing, and model optimization.
Develop predictive models using Random Forest, XGBoost, CatBoost, Gradient Boosting, Ensemble Models, Regression, and Classification algorithms.
Conduct hyperparameter tuning, cross-validation, and model evaluation using appropriate statistical and business metrics.
Deploy production-ready inference pipelines and monitor models for drift, performance degradation, and retraining requirements.
Build scalable PySpark pipelines for ingesting, cleaning, transforming, and preparing large enterprise datasets.
Develop efficient ETL/ELT pipelines supporting production ML workflows.
Optimize Spark jobs for performance and scalability.
Work with Databricks, Snowflake, Delta Lake, or similar big data platforms.
Write clean, maintainable, production-quality Python code.
Build scalable REST APIs and backend services supporting ML inference.
Participate in code reviews and follow software engineering best practices.
Build automated testing, deployment, monitoring, and retraining pipelines.
Deploy ML models into production environments and implement monitoring and alerting strategies.
Track model performance using appropriate business and technical metrics.
Collaborate with Data Engineering and Software Engineering teams to operationalize ML solutions.
Troubleshoot production issues and optimize model and system performance.
Mentor junior Machine Learning Engineers and provide technical guidance on model development and production best practices.
Required Qualifications
5+ years of hands-on Machine Learning Engineering experience.
Strong expertise in Python programming.
Strong experience writing production PySpark code.
Strong understanding of data science, statistics, and Machine Learning fundamentals.
Strong understanding of deep learning and NLP fundamentals.
Experience building and deploying production Machine Learning models.
Experience with Databricks, Snowflake, or similar large-scale data platforms.
Strong understanding of the complete ML lifecycle, including data preparation, feature engineering, model training, hyperparameter tuning, model evaluation, production deployment, monitoring, and retraining.
Experience developing scalable data pipelines and distributed data processing solutions.
Strong SQL skills.
Experience with Scikit-learn and Machine Learning algorithms including Random Forest, XGBoost, CatBoost, Gradient Boosting, Regression, and Classification.
Experience with model evaluation, feature engineering, hyperparameter tuning, cross-validation, model monitoring, and drift detection.
Experience with ETL/ELT and distributed data processing.
Experience working in Agile software development environments.
Experience writing production-quality Python code and building scalable ML solutions.
Experience deploying and monitoring Machine Learning models in production environments.
Preferred Qualifications
Experience building transformer-based recommendation models.
Familiarity with multi-armed bandit approaches.
Experience with Retrieval-Augmented Generation (RAG) solutions.
Experience with LangChain or LangGraph.
Experience integrating LLM APIs into enterprise applications.
Experience with Vector Databases.
Experience with MLOps tools such as MLflow.
Experience with cloud platforms including AWS or Azure.
Experience with Docker.
Experience with CI/CD pipelines and Git.
Experience with Generative AI, RAG, or AI agents.