SME Platform Engineer

General Dynamics Information Technology, Inc.

$191K — $258K *
Information Technology
11 - 15 years of experience
Job Overview by Ladders

Qualifications

  • 15+ years of related experience
  • US Citizenship required
  • Top Secret/SCI clearance required (or ability to obtain)
  • Deep expertise in Kubernetes ecosystem tools, GitOps methodologies, and compliance standards
  • Strong Linux server administration and command-line skills

Responsibilities

  • Architect, deploy, and manage on-premises cloud infrastructure using RKE2
  • Host and maintain robust data science environments including POSIT Workbench/Connect
  • Manage and scale GPU and Ray clusters for machine learning
  • Build and optimize CI/CD pipelines using Git, Helm charts, and ArgoCD
  • Ensure continuous FIPS compliance and manage system vulnerabilities
  • Implement and maintain authentication/authorization mechanisms using Keycloak and OPA
  • Manage container images and troubleshoot issues via command-line

Benefits

  • Variety of medical plan options and 401(k) matching
  • Paid time off plans including vacations, holidays, and parental leave
  • Short and long-term disability benefits
  • Life and accident insurance coverage
  • Full flex work weeks to encourage work/life balance
Full Job Description
Type of Requisition:
Regular

Clearance Level Must Currently Possess:
Top Secret/SCI

Clearance Level Must Be Able to Obtain:
Top Secret SCI + Polygraph

Public Trust/Other Required:
None

Job Family:
IT Infrastructure and Operations

Job Qualifications:

Skills:
CI/CD, Cloud Infrastructure, Cluster Administration, Kubernetes, Linux Server Administration
Certifications:
None
Experience:
15 + years of related experience
US Citizenship Required:
Yes

Job Description:

YOUR IMPACT

Own your opportunity to work with the largest government agency in the nation. Make an impact by advancing the Department of War's mission to keep our country safe and secure.

JOB DESCRIPTION

Iron EagleX is seeking a SME Platform Engineer to support our Engineering team in Crystal City, VA. This role will lead the design, implementation, and management of our secure, on-premises cloud infrastructure. In this role, you will be the driving force behind our advanced computing environments, ensuring the seamless orchestration of containerized applications and large-scale data science platforms. You will work at the intersection of infrastructure, security, and machine learning, managing robust compute clusters and providing foundational support for AI model training and deployment. The ideal candidate has deep expertise in Kubernetes ecosystem tools, GitOps methodologies, and strict compliance standards.

MEANINGFUL WORK AND PERSONAL IMPACT

As a SME Platform Engineer, your work will directly empower our data science and engineering teams to push the boundaries of machine learning and data analytics. By building and maintaining resilient GPU and Ray clusters, you will accelerate the fine-tuning and deployment of advanced models. Your commitment to security and compliance will ensure our critical systems remain protected against vulnerabilities, providing a safe, compliant, and highly performant foundation for the organization's most impactful technical initiatives. You will not just be managing infrastructure; you will be enabling innovation.

JOB DUTIES (INCLUDE BUT ARE NOT LIMITED TO)

  • Infrastructure & Orchestration: Architect, deploy, and manage on-premises cloud infrastructure using RKE2 and maintain storage solutions like Longhorn and Object storage.


  • Platform Enablement: Host and maintain robust data science environments, including software such as POSIT Workbench/Connect and Hive Metastore.


  • AI/ML Infrastructure: Manage and scale robust GPU clusters, Ray Clusters for fine-tuning machine learning models, and VLLM Routers for efficient model inference.


  • CI/CD & Automation: Build, maintain, and optimize CI/CD pipelines using Git, Helm charts, and ArgoCD for reliable software delivery.


  • Security & Compliance: Ensure continuous FIPS compliance across the environment. Actively manage and mitigate critical and high-level vulnerabilities.


  • Identity & Access: Implement and maintain robust authentication and authorization mechanisms using Keycloak and Open Policy Agent (OPA).


  • System Administration: Pull and manage container images from secure registries such as Harbor, Docker Hub, or Containeryard. Manage all core capabilities and troubleshoot issues effectively via the command-line console.


REQUIRED SKILLS

  • Demonstrated experience designing, deploying, administering, and troubleshooting production Kubernetes environments; hands-on experience with RKE2 or similar.


  • Strong Linux systems administration skills, including the ability to manage, diagnose, and troubleshoot infrastructure and platform services through the command line (CLI).


  • Experience implementing authentication and authorization solutions using technologies such as Keycloak, Open Policy Agent (OPA), OIDC, RBAC, or comparable identity and access management frameworks.


  • Hands-on experience with Git-based development and deployment workflows, including Helm charts, CI/CD pipelines, and GitOps practices.


  • Experience with Argo CD or similar tools for declarative, GitOps-based continuous delivery.


  • Experience managing Kubernetes storage solutions, including distributed block storage and object storage; experience with Longhorn or comparable technologies preferred.


  • Experience hosting and administering data science or analytics platforms; experience with Posit Workbench, Posit Connect, Hive Metastore, or similar technologies.


  • Experience configuring and operating systems in accordance with FIPS or comparable security and compliance requirements.


  • Demonstrated experience identifying, prioritizing, and remediating critical and high-severity system and application vulnerabilities.


  • Experience administering GPU-enabled compute environments supporting AI/ML, high-performance computing, or other compute-intensive workloads.


  • Experience supporting distributed AI/ML workloads using Ray or comparable distributed computing frameworks, including model training and fine-tuning use cases.


  • Experience deploying or supporting large language model inference and serving technologies; experience with vLLM and related routing capabilities preferred.


  • Experience pulling, managing, securing, and troubleshooting container images using private or public registries such as Harbor, Docker Hub, Container Yard, or equivalent container registry platforms.


DESIRED SKILLS

  • Hands-on experience with Infrastructure as Code (IaC) and configuration management tools such as Terraform, Ansible, or comparable technologies.


  • Proficiency in scripting or programming languages such as Python, Go, or Bash to support infrastructure automation, platform operations, and troubleshooting.


  • Advanced knowledge of Linux system administration, networking concepts, and protocols within complex or highly available infrastructure environments.


  • Experience implementing and maintaining monitoring, logging, and observability solutions using tools such as Prometheus, Grafana, or comparable platforms.


  • Familiarity with MLOps practices and the machine learning lifecycle, including model development, deployment, monitoring, versioning, and operational support.


  • Experience automating infrastructure provisioning, configuration, deployment, and operational workflows in secure or regulated environments.


  • Familiarity with performance tuning, capacity planning, and resource optimization for Kubernetes, GPU, or other compute-intensive environments.


WHAT YOU'LL NEED TO SUCCEED

  • Clearance: Current TS/SCI Clearance with current or willingness to obtain CI polygraph


  • Experience:15+ years of related experience


  • Education: Bachelor's degree in Computer Science, Software Engineering, or a related field (or equivalent experience)


  • Role requirements: Work is onsite in Crystal City, VA with optional CONUS travel


  • Due to US Government Contract Requirements, only US Citizens are eligible for this role


OWN YOUR OPPORTUNITY
Explore a career at GDIT and you'll find endless opportunities to grow alongside colleagues who share your passion for the mission and delivering results. #iexjobs #iexpriority

The likely salary range for this position is $191,250 - $258,750. This is not, however, a guarantee of compensation or salary. Rather, salary will be set based on experience, geographic location and possibly contractual requirements and could fall outside of this range.

Scheduled Weekly Hours:
40

Travel Required:
10-25%

Telecommuting Options:
Onsite

Work Location:
USA VA Arlington

Additional Work Locations:

Total Rewards at GDIT:
Our benefits package for all US-based employees includes a variety of medical plan options, some with Health Savings Accounts, dental plan options, a vision plan, and a 401(k) plan offering the ability to contribute both pre and post-tax dollars up to the IRS annual limits and receive a company match. To encourage work/life balance, GDIT offers employees full flex work weeks where possible and a variety of paid time off plans, including vacation, sick and personal time, holidays, paid parental, military, bereavement and jury duty leave. To ensure our employees are able to protect their income, other offerings such as short and long-term disability benefits, life, accidental death and dismemberment, personal accident, critical illness and business travel and accident insurance are provided or available. We regularly review our Total Rewards package to ensure our offerings are competitive and reflect what our employees have told us they value most.

Our Identity Verification Process:
As part of the hiring process, we will ask you to complete an identity verification process that leverages advanced biometrics and artificial intelligence to ensure authenticity and protect against identity fraud. You are expected to be on camera during virtual interviews. We reserve the right to take your picture to verify your identity and prevent fraud. By proceeding, you authorize the collection, processing, and use of your biometric data for identity verification and security purposes.

Similar Jobs

More Jobs at General Dynamics Information Technology, Inc.

More Information Technology Jobs

Find similar SME Platform Engineer jobs: