Application deadline: Aug 13, 2026
This position requires that the candidate selected be a US Citizen and must currently possess and maintain an active TS/SCI security clearance with polygraph.
Key job responsibilities
- Designing, developing, and implementing complex, scalable, and secure solutions tailored to customer needs
- Providing technical guidance and troubleshooting support throughout project delivery
- Integrate multiple third-party software products into a unified platform stack
- Acting as a trusted advisor to customers on industry trends and emerging technologies
- Sharing knowledge within the organization through mentoring, training, and creating reusable artifacts
BASIC QUALIFICATIONS
- 7+ years of technical specialist, design and architecture experience
- 7+ years of external or internal customer facing, complex and large scale project management experience
- 5+ years of software development with object oriented language experience
- 5+ years of continuous integration and continuous delivery (CI/CD) experience
- Current, active US Government Security Clearance of TS/SCI with Polygraph
PREFERRED QUALIFICATIONS
- Experience architecting and operating Kubernetes platforms on bare metal at scale, including lifecycle management, cluster provisioning, and day-2 operations (patching, upgrades, failure recovery) in air-gapped or high-security environments. Experience with SpectroCloud PaletteAI or similar bare-metal Kubernetes lifecycle management platforms is strongly preferred.
- Hands-on experience with NVIDIA AI Enterprise, Run:AI, or equivalent GPU orchestration and scheduling platforms for large-scale AI/ML workloads, including multi-tenant resource management, gang scheduling for distributed training, and GPU fleet health monitoring.
- Demonstrated experience integrating multiple third-party software products into a unified platform stack, including managing cross-vendor compatibility, version dependencies, and developing cohesive operational procedures across components from OS through application layer.
- Familiarity with high-performance networking in GPU cluster environments and network automation for multi-tenant isolation, including zero-trust network architectures and micro-segmentation approaches for securing multi-tenant AI infrastructure.
- Experience developing and executing incremental deployment strategies for complex platforms, including automated provisioning, infrastructure-as-code, and validation testing at progressively larger scale in environments where no prior playbook exists.
The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.
USA, CO, Denver - 153,600.00 - 207,800.00 USD annually
USA, VA, Arlington - 153,600.00 - 207,800.00 USD annually
USA, VA, Herndon - 153,600.00 - 207,800.00 USD annually
USA, WA, Seattle - 153,600.00 - 207,800.00 USD annually