Principal DevOps Engineer - Azure

Tenex.AI Inc

$150K — $180K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 8+ years in DevOps, Site Reliability Engineering, or Platform Engineering roles.
  • Expertise in building Azure platforms, including AKS and Azure Policy.
  • Experience with migrating production workloads across cloud environments.
  • Proven ability to create secure and compliant environments (SOC 2, ISO 27001).
  • Strong knowledge of microservices, containerization, and event-driven architecture.
  • Competence with Infrastructure-as-Code tools like Terraform or Bicep.
  • Production experience in Go or Python, not just scripting.

Responsibilities

  • Own architecture of Azure platform for scaling security data and events.
  • Manage Azure governance, including policy and network topology.
  • Lead SRE initiatives and enforce Service Level Objectives.
  • Implement monitoring, observability, and disaster recovery strategies.
  • Establish DevSecOps practices and CI/CD pipelines for secure delivery.
  • Automate deployment and management of microservices using containers.
  • Collaborate with engineering teams for optimization of performance and cost.

Benefits

  • Work in a collaborative in-person culture, emphasizing community.
  • Flexible work arrangement with WFH on Fridays.
  • Mentorship opportunities in Azure architecture and best practices.
Full Job Description


As a Principal DevOps Engineer, you will be a key technical leader responsible for the architecture, evolution, and operation of our Azure infrastructure, CI/CD pipelines, and Site Reliability Engineering (SRE) practices. You will keep the platform highly available, secure, and performant as it scales to handle petabytes of security data and billions of daily events.

You'll work closely with Software Engineering, AI/ML, and Security Operations teams to define the technical vision and architecture for our production systems, driving automation and operational excellence to minimize toil and accelerate product delivery. This role requires deep, hands-on Azure platform expertise, a strong software engineering foundation, and fluency in DevSecOps principles.

Culture is one of the most important things at TENEX.AI. Explore our culture deck at culture.tenex.ai to witness how we embody it, prioritizing the irreplaceable collaboration and community of in-person work.

Location: This role will require Monday through Thursday onsite in our Kansas City office (preferred), with San Jose or Sarasota, FL also considered. WFH Friday. Candidates must live in or be willing to relocate to one of these three cities.
Job Responsibilities
  • Own the architecture of our Azure platform as it scales to petabytes of security data and billions of daily events.
  • Own our Azure governance and environment model, including subscription and management group structure, Azure Policy, network topology, and identity.
  • Lead Site Reliability Engineering (SRE) initiatives, defining and driving adherence to critical Service Level Objectives (SLOs) and Service Level Indicators (SLIs), and managing on-call rotations.
  • Drive operational excellence by implementing advanced monitoring, observability (logs, metrics, tracing), automated provisioning, and disaster recovery strategies.
  • Establish and enforce DevSecOps practices, standardizing CI/CD pipelines, infrastructure-as-code (IaC), security testing, and deployment mechanisms for rapid, secure, and reliable software delivery.
  • Automate deployment, scaling, and management of microservices and event-driven systems using containerization and orchestration technologies (Docker, Kubernetes, AKS).
  • Maintain workload portability across the platform so that infrastructure decisions remain reversible.
  • Partner with engineering teams to optimize application performance, resource utilization, and cloud cost efficiency.
  • Mentor and influence engineering teams on best practices in Azure architecture, reliability, and security-first development.
  • Collaborate with Product Management and Security Operations to translate new product requirements and operational needs into scalable and cost-effective platform solutions.
  • Evaluate and drive the adoption of new infrastructure technologies and engineering methodologies to maintain a competitive advantage.
Required Skills & Qualifications
  • 8+ years of progressive experience in DevOps, Site Reliability Engineering (SRE), or Platform Engineering roles.
  • Deep, hands-on expertise building production Azure platforms, including AKS, Entra ID and workload identity federation, VNet design and Private Link, Key Vault, Azure Policy, and subscription or landing zone architecture. This is a platform engineering role rather than a Microsoft 365, Intune, or Windows administration role.
  • Experience standing up or moving production workloads across cloud environments, with the ability to describe the design, the data path, the cutover, and what broke.
  • Experience building secure and compliant (e.g., SOC 2, ISO 27001) environments.
  • Deep understanding of microservices architecture, containerization (Docker, Kubernetes), and event-driven systems.
  • Extensive experience with Infrastructure-as-Code tools (e.g., Terraform, Bicep) and CI/CD best practices.
  • Production experience writing and shipping software in Go or Python, beyond scripting and configuration.
  • Experience with monitoring and observability tools (Prometheus, Grafana, Azure Monitor, ELK stack, or similar).
  • Familiarity with real-time data pipelines and stream processing (e.g., Kafka, Event Hubs, Service Bus, Pub/Sub).
  • Proven track record of architecting, building, and operating highly scalable, distributed, and secure enterprise-grade SaaS platforms.

Nice-to-have
  • Working knowledge of more than one major cloud provider, deep enough to judge where the provider models differ rather than assume they match.
  • Prior experience in cybersecurity (SIEM, EDR, SOAR, or MDR) or an MSSP environment.
  • Experience with large-scale data warehousing/lakehouse technologies (e.g., Azure Data Explorer, Microsoft Fabric, Snowflake, BigQuery).
  • Background leading technical initiatives in high-growth startups or enterprise SaaS.
  • Familiarity with the underlying infrastructure to support AI/ML model deployment and monitoring (MLOps).
Education & Certifications
  • 10-12 years of experience, Bachelor's or Master's degree in Computer Science, Engineering, and or years of relative experience
  • Relevant certifications (Azure Solutions Architect Expert, Kubernetes, or security-related credentials) are a plus. Certifications complement production depth and do not substitute for it.

More Jobs at Tenex.AI Inc

More Information Technology Jobs

Find similar Principal DevOps Engineer - Azure jobs: