Principal Azure Platform & Cloud Operations Architect

Kastech Software Solutions Group

$125K — $150K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's or master's degree in Computer Science, Engineering, or related field
  • 8+ years of hands-on experience in cloud platform engineering, DevOps/SRE, or cloud operations
  • Strong Azure experience, especially with Azure Kubernetes Service (AKS)
  • Deep expertise in Azure networking and secure connectivity patterns
  • Proven experience with Infrastructure as Code using Terraform
  • Strong background in GitOps and deployment tools like Argo CD and Helm
  • Proficiency managing CI/CD pipelines using Azure DevOps, Jenkins, or GitHub Actions

Responsibilities

  • Assess and improve cloud operations processes, implementing automation and operational guardrails
  • Own and enhance AKS platform architecture, focusing on design and operational maturity
  • Lead Azure networking architecture, providing expertise and troubleshooting support
  • Design and implement infrastructure as code (IaC) using Terraform with best practices
  • Advance GitOps maturity with Argo CD and Helm for standardized deployments
  • Enhance CI/CD pipelines for quality gates, validation, and automated delivery
  • Implement observability practices using Azure monitoring tools to enhance operational readiness
  • Act as L3/L4 escalation point for incidents, guiding recovery and improvements
  • Mentor and coach the Cloud Operations team to elevate engineering standards
  • Collaborate with multiple teams to ensure alignment on operational practices and readiness

Benefits

  • Flexible work arrangements
  • Professional development opportunities
  • Access to the latest cloud technologies and tools
  • Collaborative and inclusive team culture
  • Opportunities for mentoring and leadership training
Full Job Description
Opportunity Overview

We are looking for a hands-on Principal Azure Platform & Cloud Operations Architect to assess

and improve how we run critical production systems on Azure. You will evaluate our current

Cloud Operations processes and platform architecture, identify automation and improvement

opportunities, implement stronger operational patterns, and act as the escalation SME when the

team hits technical roadblocks-especially across AKS, networking, and deployments.

Key Responsibilities
• Assess and Improve Cloud Operations Processes: Review current operational

workflows (provisioning, deployments, incident response, change management) and

implement process corrections, automation opportunities, and operational guardrails to

reduce manual effort and improve reliability.
• Own AKS Platform Architecture and Operations: Lead the design and operational

maturity of Azure Kubernetes Service (AKS) environments, including cluster topology,

node pools, upgrades, scaling, resiliency patterns, ingress/egress, workload identity,

secrets, and runtime security.
• Lead Azure Networking Architecture and Troubleshooting: Provide deep expertise

in Azure networking and connectivity patterns (VNET design, routing/UDRs, NSGs,

DNS, private endpoints, firewalls, load balancers, gateways, and secure egress/ingress)

and troubleshoot complex network and performance issues impacting production

systems.
• Deliver Hands-on Infrastructure as Code (IaC): Design and implement IaC using

Terraform with reusable modules, clear lifecycle management, environment consistency,

and safe change practices.
• Advance GitOps and Deployment Standardization: Strengthen deployment maturity

using Argo CD and Helm, improving repeatability, release confidence, environment

promotion, and rollback strategies.
• Improve CI/CD and Release Automation: Enhance CI/CD pipelines (Azure DevOps /

Jenkins / GitHub Actions) to implement quality gates, validation, security scanning, and

automated delivery patterns to production.
• Implement Observability and Operational Readiness: Improve monitoring, logging,

alerting, and dashboards using Azure Monitor, Log Analytics, and Application Insights to

create actionable signals and reduce noise; promote production readiness practices

(runbooks, readiness reviews, operational checklists).
• Provide L3/L4 Escalation and Incident Leadership: Act as the technical escalation

point for high-severity incidents, guiding triage and recovery, leading root cause analysis,

and ensuring corrective/preventive actions are implemented through automation and

platform improvements.
• Coach and Unblock the Cloud Operations Team: Mentor engineers and provide

hands-on guidance during complex technical challenges, raising overall capability and

establishing consistent engineering standards.
• Collaborate Across Teams: Work closely with Engineering, SRE, Security, and Delivery

teams to align operational patterns, platform guardrails, and production readiness across

services and environments.

Qualifications
• Bachelor's or master's degree in Computer Science, Engineering, or a related field.
• 8+ years of hands-on experience in cloud platform engineering, DevOps/SRE, or cloud

operations, with ownership of production-grade systems.
• Strong hands-on experience with Azure, particularly Azure Kubernetes Service (AKS),

and deep experience running Kubernetes in production (upgrades, scaling, failure

modes, troubleshooting).
• Deep expertise in Azure networking and secure connectivity patterns, with the ability to

diagnose complex multi-layer issues across AKS + network + application boundaries.
• Proven hands-on experience implementing IaC with Terraform (modules, state strategy,

environment consistency, safe rollout practices).
• Strong experience with GitOps and deployment tooling, including Argo CD and Helm,

and a strong understanding of release strategies and operational controls.
• Proficiency managing CI/CD pipelines and automation (Azure Pipelines, Jenkins, GitHub

Actions) and improving deployment reliability through automated checks and gates.
• Hands-on experience with Azure observability tooling (Azure Monitor, Log Analytics,

Application Insights) to improve service health visibility and incident response

effectiveness.
• Proficiency in scripting/automation with Python and/or Bash/PowerShell to build

operational tooling and reduce repetitive manual work.
• Strong problem-solving and communication skills, with the ability to operate calmly under

pressure and guide teams throug

Similar Jobs

More Jobs at Kastech Software Solutions Group

More Information Technology Jobs

Find similar Principal Azure Platform & Cloud Operations Architect jobs: