Job Summary
The Infrastructure Reliability Engineer will improve reliability, automation, security, and operability across enterprise infrastructure spanning cloud and on-premises/hybrid environments. The role focuses on hands-on Terraform and Ansible automation, Windows and Unix/Linux operations, observability, incident management, postmortems, and end-to-end vulnerability remediation across servers, platforms, middleware, and containerized workloads.
Key Responsibilities
• Lead remediation of infrastructure vulnerabilities across Windows, Linux, middleware, and supporting platform components.
• Drive closure of control findings, audit items, and cybersecurity remediation commitments within agreed timelines.
• Partner with Cybersecurity, Infrastructure, and Application Development teams to identify, prioritize, and remediate vulnerabilities at scale.
• Establish sustainable patching, upgrade, and lifecycle management processes to reduce recurring vulnerabilities and findings.
• Build and maintain Infrastructure as Code using Terraform across cloud and/or virtualized environments.
• Automate configuration, provisioning, patching, and deployments using Ansible across Linux/Unix and Windows environments.
• Standardize development, test, staging, and production environments through reusable Terraform modules and Ansible playbooks.
• Enforce configuration consistency and help prevent infrastructure drift.
• Operate and troubleshoot enterprise infrastructure components end-to-end, including virtual machines, servers, and virtualization platforms such as VMware or equivalent technologies.
• Support networking components including DNS, routing, VPNs, proxies, load balancers, firewalls, and other security controls.
• Partner with application teams to ensure infrastructure supports scalable and reliable application delivery.
• Own vulnerability remediation workflows across operating systems, middleware, images, and dependencies, including scanning, triage, patching/upgrading, validation, and reporting.
• Support infrastructure hardening standards, including baseline configurations, least privilege, secrets handling, and access controls.
• Support the closure of security and audit findings.
• Implement and support CI/CD automation using tools such as Jenkins, Spinnaker, and cloud deployment platforms.
• Enable safe release patterns such as blue/green deployments, canary releases, and automated rollback.
• Implement and enforce quality and security gates using tools such as SonarQube and vulnerability scanning solutions.
• Build and maintain monitoring, alerting, and dashboards using Prometheus, Grafana, Dynatrace, Splunk, and relevant cloud-native monitoring tools.
• Improve MTTR through enhanced metrics, logs, traces, service dashboards, and clearly defined escalation paths.
• Define and drive reliability outcomes related to availability, latency, resilience, and recoverability using SLIs and SLOs where applicable.
• Reduce alert fatigue by tuning alerts, improving signal-to-noise ratios, and maintaining actionable operational runbooks.
• Participate in on-call support, incident management, root-cause analysis, postmortems, and operational improvement activities.
Required Qualifications
• Minimum 5+ years of experience in SRE, DevOps, Infrastructure Engineering, Production Operations, or a related discipline.
• Proven hands-on experience implementing and supporting Terraform and Ansible automation in production environments.
• Strong administration and troubleshooting experience across Windows and Unix/Linux environments.
• Experience supporting enterprise infrastructure including compute, networking, storage, load balancing, virtualization, and security controls.
• Demonstrated ownership of vulnerability remediation, including actual patching, upgrading, validation, and reporting rather than monitoring alone.
• Experience with vulnerability management tools such as Qualys, Tenable, Rapid7, or similar platforms.
• Practical experience with incident management, on-call support, root-cause analysis, postmortems, and operational improvements.
• Experience working across Cybersecurity, Risk, Controls, Infrastructure, and Application Development teams.
• Strong understanding of infrastructure security, hardening, least privilege, secrets handling, and access controls.
• Experience with CI/CD automation and release enablement.
• Experience with observability and monitoring tools such as Prometheus, Grafana, Dynatrace, Splunk, or equivalent technologies.
• Strong analytical, troubleshooting, communication, and problem-solving skills.
Preferred Qualifications
• Experience with audit and controls remediation.
• Experience supporting cloud and hybrid/on-premises enterprise environments.
• Experience with VMware or equivalent virtualization platforms.
• Experience with Jenkins, Spinnaker, SonarQube, or similar CI/CD and quality/security tooling.
• Experience defining and applying SLIs, SLOs, reliability metrics, and operational runbooks.