Full Job Description
About the Role
The Systems Development Engineer II will own the operational health, resiliency, and automation of team infrastructure including pipelines that support integration and verification for ALG programs. This role sits at the intersection of systems engineering, software development, and operational excellence - you will define how teams build, deploy, and maintain reliable test infrastructure, not just how they test.
Working within Amazon Leo for Government (ALG) Engineering Integrated Systems Solutions, you will architect and deliver infrastructure-as-code, CI/CD pipelines, monitoring and observability solutions, and automation frameworks that keep hardware-in-the-loop (HIL) and system lab environments healthy, resilient, and scalable. You will serve as the technical owner of pipeline infrastructure reliability, driving decisions on system architecture, deployment patterns, and operational best practices while collaborating across inter-disciplinary teams, partners, and suppliers.
You will identify and eliminate risks to system resiliency, coach others on operational excellence, and use processes such as code reviews, operational readiness reviews (ORRs), and post-incident analyses (COEs) as teaching mechanisms. When systems fail, you ensure contributing causes are identified and eliminated with permanent fixes.
The successful candidate will have demonstrated experience working on complex hardware/software systems in a distributed environment, with a foundation in Linux, networking, and infrastructure automation.
Security Clearance: This position requires that the candidate selected be a US Citizen and candidates must be able to obtain and maintain a US Government security clearance of TS/SCI.
Key job responsibilities
Own the design, deployment, operational health, and long-term resiliency of lab infrastructure, test environments, and CI/CD pipelines consumed by I&T engineers and subsystem teams in hardware-in-the-loop (HIL) and system lab environments
Architect and implement infrastructure-as-code solutions, monitoring/observability platforms, and automation that ensure lab systems are reliable, scalable, and cost-effective
Identify and mitigate risks to system resiliency through participation in design reviews, ORRs, and post-incident analyses; drive permanent fixes for contributing causes of failures
Define and drive operational best practices across the lab engineering organization, establishing patterns for automation, observability, alerting, and incident response that raise the bar for the broader team
Deeply diagnose problems across the full stack - hardware, software, networking, and operating environments - and drive root-cause resolution of performance, reliability, and availability issues
Develop and maintain clear, inclusive documentation including runbooks, architecture diagrams, and operational procedures; ensure systems can be maintained and extended by those not familiar with them
Operate independently to identify gaps in infrastructure reliability, propose solutions, and drive implementation with pragmatic tradeoffs between speed and long-term maintainability
Coach and mentor team members on identifying and eliminating risk, operational excellence, and systems engineering best practices; participate actively in hiring
Collaborate with systems engineers, hardware teams, software developers, and external partners to ensure infrastructure supports all phases of integration, test, verification, and validation across satellite and ground element campaigns
Begin contributing to strategic and operational planning (OP1/OP2), ensuring the team makes appropriate investments in agility and resiliency
BASIC QUALIFICATIONS
- Experience in automating, deploying, and supporting large-scale infrastructure
- Experience programming with at least one modern language such as Python, Ruby, Golang, Java, C++, C#, Rust
- Experience with Linux/Unix
- Experience with CI/CD pipelines build processes
PREFERRED QUALIFICATIONS
- Experience with distributed systems at scale
The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.
USA, CA, NORTHRIDGE - 129,200.00 - 174,800.00 USD annually