About the RoleOtter.ai is seeking a Director of Engineering, Infrastructure to lead the teams that power our products at scale. This leader will own our cloud infrastructure, CI/CD, DevOps, site reliability practices, and production operations.
This is a hands-on, high-impact leadership role for someone who can set infrastructure strategy while going deep on critical technical and operational challenges. You will partner closely with Product, Product Engineering, AI/ML, Security, and Finance to deliver a platform that is reliable, scalable, secure, developer-friendly, and cost-efficient.
Your Impact- Own the architecture and operation of cloud, compute, networking, database, production, and AI inference infrastructure.
- Translate business and product plans into infrastructure roadmaps, capacity requirements, staffing plans, and investment decisions.
- Ensure production infrastructure is reliable, scalable, secure, performant, and cost-efficient.
- Establish infrastructure metrics, engineering standards, observability, and clear accountability, proactively surfacing risks and opportunities.
- Improve developer velocity through automation, CI/CD, and self-service tooling.
- Strengthen incident response, disaster recovery, and business continuity practices.
- Partner with AI/ML teams to optimize inference performance, reliability, resource utilization, and cost.
- Embed access control, compliance, and risk management into infrastructure architecture and operations.
- Recruit, develop, and lead a high-performing infrastructure engineering organization as a technically credible player-coach.
- Communicate infrastructure priorities, risks, investments, and tradeoffs clearly to the executive team, and partner with Engineering and Finance on cost and capacity decisions.
What We're Looking For- Bachelor's degree or higher in Computer Science or a related technical field.
- 10+ years of software development or engineering experience, including 4+ years of engineering management experience.
- Significant experience leading infrastructure, DevOps, SRE, or platform teams responsible for production systems at scale.
- Deep experience designing and operating highly available, fault-tolerant distributed systems.
- Strong knowledge of AWS and cloud-native infrastructure, including Kubernetes/EKS, EC2, Auto Scaling Groups, load balancing, networking, storage, IAM, monitoring, and infrastructure-as-code technologies such as Terraform.
- Deep understanding of MySQL operations, including performance, replication, clustering, backup, and disaster recovery.
- Strong knowledge of asynchronous architectures, including message queues, worker systems, retries, idempotency, and failure recovery.
- Experience building strong production reliability practices, including observability, incident management, on-call operations, disaster recovery, multi-region architectures, and zero-downtime migrations.
- A track record of materially improving cloud efficiency and infrastructure unit economics without compromising reliability or performance.
- Experience operating latency-sensitive, high-throughput AI inference workloads in production; GPU inference optimization experience is a plus.
- Experience scaling infrastructure to support rapid product and customer growth, including both self-serve products and enterprise customers with demanding security, reliability, and compliance requirements.
- Ability to move effectively between architecture, operational data, financial analysis, organizational design, and executive communication, with strong cross-functional influence.
- Ability to use AI-assisted engineering and operational tools to improve team productivity, incident response, diagnostics, and infrastructure management.
- Familiarity with relevant application technologies such as Python, Django, FastAPI, and WebSockets.
Salary rangeSalary Range: $244,000 to $313,000 USD per year.
This salary range represents the low and high end of the estimated salary range for this position. The actual base salary offered for the role is dependent on several factors. Our base salary is just one component of a comprehensive total rewards package.
#LI-Hybrid