Must Have Technical/Functional Skills:
• 5+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or a related role supporting cloud-based production environments.
• Practical vulnerability-management experience; familiarity with Qualys, JFrog Xray,
SCA/SBOM, or equivalent tooling; as well as knowledge of dependency-management for Ruby/Bundler, Python/pip, and Go modules. Experience with direct versus transitive dependencies, version constraints, lockfiles, go.mod/go.sum, and verifying the effective version shipped in a built artifact.
• Experience developing and maintaining infrastructure automation using Ansible.
• Experience administering and troubleshooting Linux-based systems and distributed infrastructure environments.
• Experience designing, implementing, and maintaining CI/CD pipelines, including GitLab CI.
• Experience supporting large-scale infrastructure environments consisting of hundreds or thousands of systems.
• Available to be online from 9am-4pm PST.
Roles & Responsibilities:
Own findings for FedRAMP Vulnerability Management from intake through validated remediation and closure. Triage and prioritize host, OS-package, container, base-image, and application-dependency findings using exploitability, asset criticality, and remediation timelines. Identify the real source of vulnerable software and determine whether the right action is a dependency update, image rebuild, promotion, host change, deviation, false-positive correction, or ticket cleanup. Implement or coordinate fixes through Ansible, GitLab CI, package managers, container builds, Artifactory, and Federal promotion-train workflows and validated promoted artifacts. Validate fixes using effective package versions, build artifacts, image manifests, deployed-host evidence, and scanner rescans before resolving vulnerability tickets. Maintain Jira evidence, ownership, due-date escalation, exception rationale, backlog metrics, runbooks, automation, and knowledge transfer. Success will mean reducing the overdue backlog, improving ownership and evidence quality, delivering repeatable automation and runbooks, and transferring a sustainable process to the CSI team for future use.Develop and maintain automation solutions that improve the reliability, scalability, security and operational efficiency of infrastructure and processes for the hosts in our FedRAMP cloud environments, as well as new innovations that allow us to adjust to changing compliance requirements. Design and enhance deployment pipelines, testing frameworks, and operational tooling to support the continued growth of a platform serving a rapidly growing number of managed devices worldwide. Troubleshoot complex infrastructure and distributed systems issues to ensure high availability while helping teams adapt their infrastructure, applications and processes to FedRAMP controls and requirements. Contribute to critical projects such as refactoring our deployment process, and new cluster build outs by building automation that enables manual toil reduction, as well as rapid and repeatable processes for new purpose-built cloud environments. Partner with other engineering teams, product management, and business partners across multiple teams and time zones to understand platform dependencies, seek opportunities for improvement, and deliver solu tions that enhance reliability and performance while reducing operational overhead.
Good to Have:
• Familiarity with AWS or other public cloud platforms and hybrid infrastructure environments.
• Knowledge of monitoring, observability, and reliability engineering practices and tooling.
• Familiarity with Kubernetes concepts and containerized application platforms.
• Experience employing AI-assisted development tools to improve software development, automation, operational analysis, and engineering productivity.
• Experience managing fleet wide software deployments
• Experience providing incident support and triage.
Salary Range: $86,000 - $120,000 a year
#LI-CM2