Role description
Job Description :
6 years of experience in infrastructure support site reliability engineering cloud operations or platform engineering including strong hands-on ownership of production Azure environments
Demonstrated expertise in reliability performance capacity cost optimization and high severity incident leadership
Deep knowledge of incident problem change security privacy compliance audit and governance processes
Strong experience with Azure access control certificates secrets monitoring logging CICD Infrastructure as Code and scripting
Solid understanding of data platform architecture and dependencies across ADF ADLS Gen2 Synapse Cosmos DB Azure Data Explorer SQL Server and Microsoft Fabric
Strong production ownership customer facing communication and the ability to drive operational discipline across onsite and offshore teams
Key Responsibilities
Infrastructure reliability and optimization Own the availability performance capacity and cost of production Azure infrastructure supporting high volume data platforms target 9999 availability SLA adherence and QoS while proactively addressing bottlenecks saturation and scaling risks
Azure architecture security and governance Design and operate landing zones subscriptions resource groups VNets peering ExpressRoute Azure Firewall Bastion DDoS protection Azure Policy Entra ID RBAC Managed Identities PIM and Conditional Access
Incident problem and change management Lead triage mitigation stakeholder communication root cause analysis corrective actions risk assessment approvals validation and rollback planning improve MTTR and prevent repeat incidents
Compliance and operational readiness Maintain S360 security privacy audit and governance compliance sustain accurate runbooks SOPs CENs and operational playbooks
Certificates secrets and dependencies Manage the lifecycle of certificates keys secrets identities and service dependencies track expirations automate renewals and secure service to service communication
Data platform infrastructure Support and optimize Azure Data Factory ADLS Gen2 Synapse Analytics Cosmos DB Azure Data Explorer SQL Server and Microsoft Fabric troubleshoot throughput and dependency issues and guide platform modernization and Fabric migration
Monitoring and observability Use Azure Monitor Log Analytics Application Insights and KQL platform metrics ing and cost dashboards to detect risks analyze trends and drive evidence based decisions
Automation and engineering practices Standardize infrastructure and operational workflows through Azure DevOps GitHub Actions ARM Bicep Terraform YAML PowerShell Azure CLI Python and Power Automate apply SRE practices including error budgets automated recovery and self-healing
Cost and performance management Right size resources optimize compute storage networking SQL and Cosmos capacity and implement budgets and reservation planning without compromising reliability
Customer and team collaboration Serve as the infrastructure SRE contact for Redmond customers communicate risks and optimization opportunities clearly and coordinate consistent execution across onsite offshore infrastructure SRE and data engineering teams