Senior AI Platform Engineer

Peraton

• $146K — $234K *
Technical Services
11 - 15 years of experience
Job Overview by Ladders

Qualifications

  • 12+ years of experience with a BS/BA, 10+ years with an MS/MA, or 7+ years with a Ph.D.
  • Hands-on experience with Kubernetes-based container platforms.
  • Active DoD Secret clearance required.
  • 12+ years of platform/systems/integration engineering experience.
  • Practical AI model deployment and operation skills with a focus on performance tuning.
  • Infrastructure-as-code expertise using Terraform and Ansible.
  • Experience deploying and managing OpenShift on bare metal.

Responsibilities

  • Deploy and manage large language models and integrate AI services.
  • Optimize model performance and configurations for efficiency.
  • Monitor and troubleshoot model-serving environments.
  • Transition manual installs to IaC-managed provisioning.
  • Install and operate bare-metal OpenShift clusters.
  • Design infrastructure provisioning using Terraform and Ansible.
  • Support OCONUS hardware installation and configuration.

Benefits

  • Flexible telework options available within specified locations.
  • Opportunities for professional development and certifications.
  • Engagement in cutting-edge AI technologies and platforms.
  • Support for career growth within a defense-focused environment.
Full Job Description
Responsibilities

The Senior AI Platform Engineer will stand up and operate a bare-metal container platform, support a feasibility assessment of all platform technologies, and drive the transition towards infrastructure-as-code managed provisioning using Terraform and Ansible. This individual will work in a bare-metal-first environment while actively reducing reliance on manual installs over time, and will support core platform services and OCONUS bootstrap activities.

 

This role is flexible for occasional telework however, candidates must be within 50 miles to one of the following work locations:  Oklahoma City, OK, Montgomery, AL, Ogden, UT, San Antonio, TX, Mechanicsburg, PA, Chambersburg, PA, Arlington, VA, Ft. Meade, MD, Columbus, OH.

 

Other Responsibilities Include:

  • Deploy, configure, and operate large language models (LLMs), and integrate approved frontier AI models and services into enterprise environments
  • Optimize model inference and deployment configurations to improve response times, throughput, and GPU/memory utilization while maintaining output quality
  • Monitor, troubleshoot, and maintain model-serving environments, including model updates, API connectivity, and supporting infrastructure
  • Transition platform provisioning from manual/bare-metal installs toward a fully IaC-managed model
  • Install, configure, and operate bare-metal OpenShift (or equivalent enterprise Kubernetes distribution) clusters
  • Design and implement infrastructure-as-code provisioning using Terraform and Ansible for platform and VM orchestration
  • Support the objective feasibility assessment of OpenShift versus Spectro Cloud, documenting licensing tradeoffs, operability, and long-term sustainability considerations
  • Deploy and support core platform services including authoritative DNS (e.g., PowerDNS)
  • Implement and enforce STIGs and cybersecurity compliance requirements on platform infrastructure
  • Monitor platform health, performance, and capacity using observability tools; manage platform logging and alerting
  • Support OCONUS hardware bootstrap activities including physical hardware installation and configuration
  • Document platform architecture, configuration decisions, and operational procedures
Qualifications

Required Qualifications:

  • Minimum of 12 years with BS/BA; Minimum of 10 years with MS/MA; Minimum of 7 years with Ph.D.
  • Experience with platform/systems/integration engineering 
  • Hands-on Kubernetes-based container platform experience
  • US Citizenship
  • Active DoD Secret clearance
  • DoD 8570/8140 compliant certification
  • Practical experience deploying and operating AI models, with a foundational understanding of inference and performance tuning
  • Hands-on experience deploying and administering OpenShift (or another enterprise Kubernetes distribution) on bare metal
  • Infrastructure-as-code proficiency, specifically Terraform and Ansible, for provisioning and configuration management
  • Working knowledge of virtualization/cloud orchestration platforms (Spectro Cloud, VMware, or equivalent)
  • Understanding of container platform architecture: networking (CNI), storage (CSI), cluster security, and lifecycle management
  • Ability to document platform architecture, licensing tradeoffs, and operational considerations
  • Strong Linux administration skills (RHEL)
  • Must have regular/recurring physical access to a DISA SIPR-accredited facility
  • Working understanding of agentic platform modeling/hosting/services

Preferred Qualifications:

 

  • Excellent verbal and written communication skills
  • Excellent writing and grammatical skills
  • Excellent organizational skills and attention to detail
  • Certified Kubernetes Administrator (CKA) or Certified Kubernetes Application Developer (CKAD)
  • HashiCorp Certified: Terraform Associate
  • Experience with GitOps workflows (e.g., ArgoCD, Flux) for platform configuration management
  • Experience bootstrapping platforms in disconnected or OCONUS environments including physical hardware installation
Target Salary Range$146,000 - $234,000. This represents the typical salary range for this position. Salary is determined by various factors, including but not limited to, the scope and responsibilities of the position, the individual’s experience, education, knowledge, skills, and competencies, as well as geographic location and business and contract considerations. Depending on the position, employees may be eligible for overtime, shift differential, and a discretionary bonus in addition to base pay.

Similar Jobs

More Jobs at Peraton

More Technical Services Jobs

Find similar Senior AI Platform Engineer jobs: