DevOps & AI/ML Infrastructure EngineerThe DevOps Engineer is responsible for supporting and improving cloud and ML/AI infrastructure, automating deployments, and maintaining CI/CD pipelines to ensure efficient, secure, and scalable development workflows. This role plays a crucial part in infrastructure automation, monitoring, and cloud security while collaborating with Software and ML Engineers, Product support, QA, and Security teams.
As a key member of the DevOps team, the DevOps Engineer helps manage cloud environments, CI/CD pipelines, and Infrastructure as Code (IaC), ensuring high availability and compliance with security best practices.
In this role, you'll get to:Cloud Infrastructure, Security & Reliability- Support and maintain scalable, highly available, and secure cloud infrastructure in accordance with company policies and standards.
- Provision and manage cloud resources using Infrastructure as Code (Terraform, Terragrunt, CloudFormation).
- Implement cloud security best practices, including IAM/role-based access controls, encryption, vulnerability management, and secure infrastructure configurations.
- Support containerized environments and orchestration platforms.
- Apply DevSecOps principles across infrastructure and deployment workflows.
- Participate in disaster recovery planning, testing, and recovery activities.
CI/CD, Automation & Deployment- Maintain and optimize CI/CD pipelines using tools such as GitLab CI/CD and Jenkins, supporting application and ML model deployments.
- Improve deployment reliability and support zero-downtime deployment strategies.
- Automate configuration management, infrastructure provisioning, and routine operational processes.
- Troubleshoot deployment and pipeline issues and implement improvements to prevent recurrence.
- Develop scripts and automation to reduce manual work and improve engineering efficiency.
AI & Agentic Infrastructure- Help design, deploy, operate, and secure infrastructure supporting AI and agentic products, including MCP, agents, integrations, internal tooling, and customer-facing use cases.
- Use AI-assisted engineering tools, coding copilots, and AI-driven troubleshooting to improve DevOps productivity and reduce repetitive operational work.
- Evaluate and adopt practical AI-enabled workflows that improve infrastructure management, troubleshooting, and operational efficiency.
MLOps & ML Platform Infrastructure- Operate and scale ML platform infrastructure, including Databricks interactive clusters, jobs compute, ML pipelines, and Model Serving endpoints.
- Manage production model-serving infrastructure, including compute capacity, provisioned throughput, and autoscaling for high-throughput inference workloads.
- Maintain infrastructure-level monitoring for model drift, data quality, inference performance, and serving health, while partnering with ML Engineering on model evaluation, quality thresholds, and model correctness.
- Partner with ML Engineering to support reliable CI/CD and production deployment of ML models.
Observability, Incident Response & Engineering Collaboration- Maintain monitoring, logging, metrics, and alerting solutions using tools such as Prometheus, Grafana, Coralogix, and CloudWatch.
- Support incident response and perform Root Cause Analysis (RCA) for infrastructure and deployment-related issues.
- Improve system observability through effective log aggregation, metrics collection, monitoring, and alerting.
- Partner with Software Engineers, ML Engineers, QA, and Software Engineers in Test to improve deployment workflows and integrate automated testing into CI/CD pipelines.
- Collaborate with IT Security to maintain secure cloud operations and infrastructure policies.
- Respond to engineering and Product Support requests in a timely manner and provide technical infrastructure support when needed.
- Maintain accurate internal technical and operational documentation.
- Collaborate effectively with international teams across multiple time zones
Who you are and what you'll need for this position:- 3+ years of experience in DevOps, Cloud Engineering, Site Reliability Engineering (SRE), or a similar infrastructure-focused role.
- 2+ years of hands-on experience with AWS services such as EC2, S3, RDS, Lambda, IAM, VPC, SQS, API Gateway, or similar services.
- 2+ years of experience working with containerized environments and orchestration platforms such as Kubernetes and Amazon EKS.
- Strong experience building and maintaining CI/CD pipelines using tools such as GitLab CI/CD or Jenkins.
- Hands-on experience with Infrastructure as Code using Terraform, Terragrunt, CloudFormation, or similar technologies.
- Strong Linux system administration and troubleshooting skills.
- Solid understanding of networking fundamentals, including routing, load balancing, network security, and related concepts.
- Scripting experience with Python, Bash, or similar languages to automate infrastructure and operational tasks.
- Hands-on experience using AI tools to improve engineering workflows, automation, troubleshooting, or agentic use cases.
- Experience supporting data, ML, or other compute-intensive production workloads.
- Experience with Google Cloud would be valuable, particularly for candidates who have worked across multi-cloud environments.
- Familiarity with Helm and service mesh technologies such as Istio, Linkerd, Traefik, or similar tools would be beneficial.
- Experience with serverless and event-driven architectures using technologies such as AWS Lambda, API Gateway, and SQS is a plus.
- Exposure to cloud and infrastructure security practices, including vulnerability management and tools such as Nessus, Prowler, Trivy, firewalls, or similar technologies, would be valuable.
- Knowledge of security standards, compliance requirements, and cloud security best practices is beneficial.
- Experience with observability, log analysis, and monitoring platforms such as Coralogix, Prometheus, Grafana, or similar solutions is a plus.
- FinOps experience, including cloud cost monitoring, optimization, and accountability practices, would be valuable.
- Experience with API gateways or API management platforms such as Kong, Apigee, or similar technologies is beneficial.
- Experience with MLOps platforms and practices-particularly Databricks, model serving, ML pipelines, and model monitoring-would be an advantage.
Confidence can sometimes hold us back from applying for a job. But we'll let you in on a secret: there's no such thing as a 'perfect' candidate. Have 50% of the criteria? Excited about this opportunity? Passionate about what we do at CreatorIQ? Please apply! CreatorIQ is a place where everyone can grow.
What you will get from us:- People: Work with talented, collaborative, and friendly people who love what they do.
- Guidance: Utilize our learning platform to fully get the training and tools you'll need to become successful here from your first day with us.
- Work/life harmony: 15 days of vacation, floating and company holidays, wellness benefits, and paid parental leave.
- Whole Health Package: Comprehensive medical, dental, vision, life, and disability insurance, plus additional wellness benefits.
- Planning for the future: A 401(k) plan to help you plan ahead.
- Work from home stipend: To assist you in setting up a home office that works for you.