Application deadline: Aug 3, 2026
This position requires that the candidate selected must currently possess and maintain an active TS/SCI Security Clearance with Polygraph. The position further requires the candidate to opt into a commensurate clearance for each government agency for which they perform AWS work.
Key job responsibilities
- Designing, developing, and implementing complex, scalable, and secure solutions tailored to customer needs
- Providing technical guidance and troubleshooting support throughout project delivery
- Integrate multiple third-party software products into a unified platform stack
- Act as a trusted advisor to customers on industry trends and emerging technologies
- Sharing knowledge within the organization through mentoring, training, and creating reusable artifacts
BASIC QUALIFICATIONS
- 3+ years of programming in Python, Ruby, Go, Swift, Java, .Net, C++ or similar object oriented language experience
- Experience with CloudFormation, Chef, Puppet, Salt, or Ansible in production environments
- 3+ years of cloud architecture and solution implementation experience
- Experience with continuous integration and continuous delivery
- Current, active US Government Security Clearance of TS/SCI with Polygraph
PREFERRED QUALIFICATIONS
- Experience architecting and operating Kubernetes platforms on bare metal at scale, including lifecycle management, cluster provisioning, and day-2 operations (patching, upgrades, failure recovery) in air-gapped or high-security environments. Experience with SpectroCloud PaletteAI or similar bare-metal Kubernetes lifecycle management platforms is strongly preferred.
- Hands-on experience with NVIDIA AI Enterprise, Run:AI, or equivalent GPU orchestration and scheduling platforms for large-scale AI/ML workloads, including multi-tenant resource management, gang scheduling for distributed training, and GPU fleet health monitoring.
- Demonstrated experience integrating multiple third-party software products into a unified platform stack, including managing cross-vendor compatibility, version dependencies, and developing cohesive operational procedures across components from OS through application layer.
- Familiarity with high-performance networking in GPU cluster environments and network automation for multi-tenant isolation, including zero-trust network architectures and micro-segmentation approaches for securing multi-tenant AI infrastructure.
- Experience developing and executing incremental deployment strategies for complex platforms, including automated provisioning, infrastructure-as-code, and validation testing at progressively larger scale in environments where no prior playbook exists.
The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.
USA, CO, Denver - 131,300.00 - 177,600.00 USD annually
USA, VA, Arlington - 131,300.00 - 177,600.00 USD annually
USA, VA, Herndon - 131,300.00 - 177,600.00 USD annually
USA, WA, Seattle - 131,300.00 - 177,600.00 USD annually