Applicants must be authorized to work for ANY employer in the U.S. We are unable to sponsor or take over sponsorship of an employment Visa at this time.
Principal DevSecOps & AI Platform EngineerJob SummaryWe are seeking a highly experienced Senior/Principal Cloud & AI Platform Engineer to design, build, secure, and operate scalable cloud platforms and AI/ML solutions. This role combines software engineering, cloud infrastructure, DevSecOps, SRE, security engineering, and AI/ML expertise.
The ideal candidate will have strong hands-on experience with Python, TypeScript, AWS, Terraform, CI/CD, LLMs, and RAG architectures, along with a proven ability to improve reliability, observability, security, and operational processes. This individual will also provide technical leadership and collaborate across engineering, security, infrastructure, and product teams.
Key Responsibilities- Design, develop, and maintain scalable cloud-native applications and platforms using Python and TypeScript.
- Architect and implement solutions within AWS, following cloud security, scalability, reliability, and performance best practices.
- Build and maintain infrastructure using Terraform and Infrastructure as Code (IaC) principles.
- Design, implement, and improve CI/CD pipelines to automate software delivery, testing, infrastructure deployment, and security checks.
- Apply DevSecOps practices by integrating security throughout the software development and deployment lifecycle.
- Partner with security teams to identify vulnerabilities, implement security controls, and improve cloud and application security.
- Apply SRE principles to improve system reliability, availability, scalability, and operational efficiency.
- Develop and support AI/ML solutions, including applications leveraging LLMs and Retrieval-Augmented Generation (RAG).
- Design AI-powered applications that integrate LLMs with enterprise data, APIs, and other systems.
- Implement observability solutions covering metrics, logs, traces, application performance, and infrastructure health.
- Participate in incident management, including troubleshooting production issues, root cause analysis, remediation, and prevention.
- Establish and improve monitoring, alerting, incident response, and operational procedures.
- Collaborate with engineering, DevOps, security, data, and business stakeholders to deliver enterprise solutions.
- Provide technical leadership, mentoring, and guidance to engineers and development teams.
- Identify opportunities to automate manual processes and improve engineering productivity.
- Establish engineering standards and best practices around cloud architecture, security, reliability, and AI/ML development.
Required Qualifications- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field.
- 7+ years of experience in software engineering, cloud engineering, DevOps, SRE, or a related technical discipline.
- Strong hands-on experience with Python.
- Experience developing applications or services using TypeScript.
- Strong experience with AWS cloud services and architecture.
- Hands-on experience with Terraform and Infrastructure as Code.
- Experience designing and maintaining CI/CD pipelines.
- Strong understanding of DevSecOps practices and cloud/application security.
- Experience with SRE principles, production operations, and reliability engineering.
- Experience with incident management, troubleshooting, and root cause analysis.
- Strong understanding of observability, including monitoring, logging, metrics, tracing, and alerting.
- Experience with AI/ML technologies, with practical experience working with LLMs and/or RAG architectures.