Role
Senior Platform Engineer – Azure, Data & AI
Location: United States (Remote)
Reports to: AI Lead
Employment Type: Full-time, Exempt
Direct Reports: None
What is the role?
The Senior Platform Engineer builds and runs the Azure platform that Howden's US AI and data workloads sit on. This is a hands-on engineering role: you will design, code, and operate the landing zones, the data platform, the AI runtime services, and the networking, identity, and security controls that let engineering teams ship AI capability safely and quickly.
The scope spans three connected layers. The Azure foundation: subscriptions, landing zones, networking, private connectivity, identity, and policy, all expressed as Terraform. The data platform: Databricks, Lakehouse storage, Unity Catalog, ingestion and orchestration, and the governed data products that AI systems consume. The AI platform: Azure AI Foundry, the AI Gateway, agent runtime and hosting, model deployment and quota management, and the control surfaces such as Microsoft Agent 365 and the agent registry that make agent estates visible and governable.
You will work from architecture direction set by the Leads, Group Architecture, and InfoSec, and you will own the implementation end to end. Expect to spend most of your time in Terraform, in pipelines, and in Azure, with a heavy emphasis on agent-assisted development: using coding agents and AI development tooling to move faster than a conventional infrastructure pace and knowing when to trust the output and when to rewrite it.
This role is self-directed. You will be handed an outcome and a rough shape, and you are expected to break it down, sequence it, unblock yourself, and deliver production-quality work collaborating with other team members.
What success looks like
- Azure environments for AI and data workloads provisioned entirely from code, repeatable across dev, non-production, and production, with no manual portal steps.
- A Databricks and Lakehouse platform that engineering and analytics teams use daily, with governed access, documented data products, and predictable cost.
- Azure AI Foundry, AI Gateway, and agent hosting running as a shared platform service with quota, routing, cost attribution, and usage telemetry that teams can self-serve against.
- Network, identity, and data boundaries implemented to Howden and Group InfoSec standards, evidenced rather than asserted, and accepted through architecture and security review without rework.
- Platform golden paths that let a product team stand up a compliant AI workload in days rather than weeks, using published modules, pipelines, and reference implementations.
- Agent-assisted development practices used daily and shared with the team as working examples: module scaffolding, policy generation, test harnesses, and repository conventions.
- Work delivered from a stated outcome with minimal direction, escalating early when something genuinely needs a decision above your level.
- Design input that changes outcomes: options, trade-offs, and working spikes brought to the Leads and architecture before decisions are locked.
What will you be doing?
Azure Platform Engineering
- Design, build, and operate Azure landing zones for AI and data workloads: subscription and management group structure, naming and tagging standards, Azure Policy, RBAC models, and cost boundaries.
- Provision and run the compute and hosting layer for AI services: Azure Container Apps, AKS, App Service, Azure Functions, and container registries, with sensible scaling, resilience, and resource configuration.
- Build shared platform services that product teams consume: API Management, Key Vault, Service Bus, Event Hubs, Storage, Cosmos DB, and managed identity patterns.
- Own capacity, quota, and region strategy for AI and data services, including data residency and data zone constraints.
- Keep environments reproducible. Anything created by hand in the portal gets replaced by code.
AI Platform & Agent Runtime
- Build and operate Azure AI Foundry as a shared platform capability: resource and project topology per environment and data zone, model deployments, quota and throughput management, content safety configuration, and connection management.
- Implement and run the AI Gateway layer using Azure API Management AI Gateway and Foundry control plane capabilities: model routing, token-based rate limiting, semantic caching, cost attribution by team and workload, key and identity management, and usage telemetry.
- Stand up and operate the runtime that hosts AI agents: container and serverless hosting, orchestration runtimes, tool and MCP server connectivity, secrets and credential handling, and environment promotion.
- Implement agent identity and control surfaces including Microsoft Agent 365, Entra Agent ID, and the agent registry: registration, ownership, entitlement, lifecycle, and decommissioning of agents in the estate.
- Integrate the platform with governance and service management tooling so that agent inventory, approvals, and change records stay current rather than being maintained by hand.
- Build the guardrails that make self-service safe: landing-zone templates, workload onboarding automation, quotas, policy checks, and default observability wired in from day one.
Data Platform Engineering
- Build and operate the Databricks platform: workspace topology, clusters and serverless compute, Unity Catalog, cluster policies, job orchestration, and cost controls.
- Implement Lakehouse storage and medallion layering on ADLS Gen2 with Delta, including partitioning, retention, and lifecycle management.
- Build ingestion and transformation pipelines from insurance source systems using Data Factory, Databricks workflows, Microsoft Fabric, or event-driven patterns, with lineage, versioning, change detection, and reconciliation.
- Implement data governance in the platform: catalog and metadata management, classification, data quality checks, access controls, masking, and audit trails, working with Purview or equivalent tooling.
- Expose governed data and retrieval layers that AI systems depend on, including vector stores and search indexes, and keep them current through automated refresh.
- Write the SQL, Python, and transformations needed to answer your own data questions rather than waiting on another team.
Security, Identity & Networking
- Design and implement Azure network architecture for AI and data workloads: hub-and-spoke topology, VNets and subnets, private endpoints and Private Link, DNS, firewalls, NSGs, WAF, and egress control.
- Implement identity and access end to end: Entra ID, managed identities, service principals, workload identity federation, OAuth 2.0 flows, conditional access, and least-privilege RBAC across Azure, Databricks, and AI services.
- Secure secrets, keys, and certificates through Key Vault with rotation and automated distribution to workloads.
- Implement data protection controls: encryption at rest and in transit, customer-managed keys where required, PII handling, redaction in AI pipelines, and network isolation of model and data endpoints.
- Work with InfoSec and Group Architecture on threat modelling, security review, vulnerability and patch management, and remediation of findings against platform components.
- Build compliance evidence into the platform: policy-as-code, drift detection, configuration baselines, and audit logging that stand up to internal and external review.
Infrastructure as Code & Deployment Automation
- Write and maintain Terraform for the full estate: Foundry and AI services, Databricks, networking, identity and role assignments, Key Vault, storage, monitoring, and policy.
- Own module structure, state management, workspace and environment separation, variable and secret handling, versioning, and drift detection for the infrastructure you build.
- Build CI/CD pipelines in Azure DevOps or GitHub Actions covering plan and apply gates, automated testing, security and policy scanning, container builds, environment promotion, and rollback.
- Publish reusable modules and reference implementations with documentation so other teams can consume the platform without asking you first.
- Automate onboarding of new workloads and environments, so provisioning is a pipeline run, not a project.
Agent-Assisted Development
- Use coding agents and AI development tooling (Claude Code, GitHub Copilot, and similar) as a primary part of your daily workflow for implementation, refactoring, test generation, and debugging.
- Build and maintain the scaffolding that makes agent-assisted development work on our repositories: context files, tool definitions, repository conventions, task decomposition patterns, and reusable prompt assets.
- Apply engineering judgment to agent output. Review generated infrastructure and code as rigorously as human-written work and know where the tooling saves hours and where it creates cleanup work.
- Automate repetitive platform works with scripted age