Infrastructure Engineer

Dexian

$125K — $150K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in computer science or related field; advanced degrees preferred.
  • 10+ years of experience in infrastructure capacity and performance engineering.
  • Experience in regulated environments with high-assurance standards.
  • Strong data analysis skills for forecasting and actionable insights.
  • Proven process discipline in change and configuration management.
  • Experience with diverse engineering teams and cross-platform alignment.
  • Proficiency in Azure services, including identity management and security policies.

Responsibilities

  • Own the end-to-end Capacity Management model for Azure services on high-criticality projects.
  • Ensure capacity meets SLAs, RTOs, and RPOs, with ongoing regional monitoring.
  • Partner with site reliability engineers to apply capacity practices and resilience patterns.
  • Contribute to documentation and compliance evidence for continuous monitoring.
  • Develop service-level capacity models and establish buffer standards based on criticality.
  • Design autoscaling policies to maintain performance stability and manage quotas.
  • Conduct analyses on utilization and performance metrics to inform tuning and architectural decisions.

Benefits

  • Opportunity to lead critical cloud initiatives in a high-regulation environment.
  • Collaboration with cutting-edge tools and methodologies.
  • Professional growth through strategic planning and cross-functional collaboration.
  • Chance to contribute directly to complex, high-stakes projects.
  • Focus on continuous improvement in cloud operations.
Full Job Description
Position Summary

We are seeking an experienced Infrastructure Engineer with expertise in cloud capacity management to join our team. This role focuses on leading capacity planning and optimization for a high-demand Azure public cloud environment. The ideal candidate will ensure the availability of resilient, scalable capacity across compute, storage, network, and platform services. They will collaborate with site reliability teams to meet performance and reliability targets, while maintaining compliance with rigorous program controls. This position offers an opportunity to shape critical cloud infrastructure capabilities in a regulated environment, emphasizing evidence-based practices and continuous monitoring.
Key Responsibilities
  • Own and manage the end-to-end Capacity Management operating model for Azure services involved in high-criticality projects, including planning, modeling, forecasting, monitoring, tuning, and governance.
  • Ensure sufficient capacity and engineered buffers to meet service-level agreements (SLAs), recovery time objectives (RTOs), recovery point objectives (RPOs), and contractual or regulatory requirements, with particular attention to region-specific restrictions and ongoing monitoring.
  • Partner with site reliability engineers to implement capacity practices via infrastructure as code (IaC), gated change controls, performance baselines, autoscaling strategies, and resilience patterns.
  • Contribute to documentation and compliance evidence such as system security plans, control narratives, corrective action plans, and continuous monitoring artifacts.
  • Develop and maintain service-level capacity models across various Azure components, establishing buffer standards based on service criticality and validating against demand patterns and failover scenarios.
  • Design and tune autoscaling policies, setting guardrails on quotas and throttling to ensure performance stability.
  • Conduct baseline and trend analyses of utilization, throughput, and performance metrics, translating insights into tuning actions, reservations, savings plans, and architectural improvements.
  • Forecast future demand based on product roadmaps and business growth, translating forecasts into capacity plans and procurement strategies.
  • Participate in change review processes, ensuring capacity and security impacts are properly assessed and documented.
  • Manage cryptographic mechanisms and cryptography-related configurations under change control, maintaining versioned inventories and validation compliance.
  • Oversee external services supporting capacity, confirming they meet required standards and conduct ongoing oversight.
  • Enforce region-restriction policies for processing, storage, backups, and disaster recovery specific to high-impact systems.
  • Balance performance, resilience, and cost-efficiency through resource rightsizing, tiering, and scheduled scaling, ensuring proactive capacity adjustments.
  • Perform criticality analysis to prioritize capacity needs, aligning backup, monitoring, and security policies accordingly.
  • Validate disaster recovery (DR) capacity and ensure buffers are maintained for failover scenarios without impacting steady-state operations.
  • Define, measure, and report key capacity KPIs, including utilization, saturation, headroom, runway duration, scaling effectiveness, quota use, DR readiness, and cost-performance metrics.
  • Prepare dashboards and reports to monitor program compliance, support audits, and inform strategic decision-making.
Required Qualifications
  • Bachelor's degree in computer science or a related technical field; advanced degrees are a plus.
  • 10+ years of experience in infrastructure capacity and performance engineering across compute, storage, network, and platform services.
  • Demonstrated experience working within regulated environments and familiarity with high-assurance concepts and evidence collection standards.
  • Strong data analysis skills, with the ability to interpret telemetry and forecast data into actionable insights.
  • Proven process discipline with change and configuration management, and the ability to manage dependencies across IT operations and finance teams.
  • Experience coordinating diverse engineering teams and aligning delivery across multiple platforms and tools.
  • Proficiency with Azure services and concepts such as identity management, SQL, storage, networking, security policies, and role-based access control.
  • Excellent communication skills with the ability to manage stakeholder relationships and present technical information to executive audiences.
  • Capacity to translate complex technical requirements into clear plans, milestones, and measurable outcomes.
Preferred Qualifications
  • Experience developing capacity models for multi-region architectures with strict geographic restrictions.
  • Proven success in disaster recovery planning and execution, validated failover capacity, and documented evidence.
  • Experience managing corrective action and continuous monitoring submissions within regulated environments.
  • Strong collaboration with site reliability teams on service level objectives, error budgets, and reliability engineering patterns.
Why This Opportunity May Be IfAppeling

This position provides the chance to lead critical cloud infrastructure initiatives within a high-regulation environment, influencing how capacity and resilience are managed at an enterprise scale. You will work with cutting-edge tools and methodologies to ensure service availability and compliance, contributing directly to the success of complex, high-stakes projects. The role supports professional growth through involvement in strategic planning, cross-functional collaboration, and continuous improvement in cloud operations.

Similar Jobs

More Jobs at Dexian

  • Infrastructure Engineer
    $125K — $150K *
    New York, NY 10025 (New York County)
    Information Technology
    In-Person
  • Financial Analyst
    $95K — $115K *
    New York, NY 10025 (New York County)
    Finance & Insurance
    In-Person
  • IT Project Manager (JDE1)
    $95K — $115K *
    Fort Washington, PA 19034 (Montgomery County)
    Information Technology
    In-Person

More Information Technology Jobs

Find similar Infrastructure Engineer jobs: