Platform Operations Engineer - TS/SCI

Sunayu, LLC

$110K — $130K *
Aerospace & Defense
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • BS in Engineering, Computer Science, or related field with 8+ years of experience; 6+ years with a Master's; additional experience may substitute for a degree
  • Active TS/SCI clearance with polygraph eligibility
  • At least one DoD 8570.01 M IAT Level II+ certification
  • Experience with Kubernetes and enterprise scale production systems
  • Strong communication skills and ability to perform under pressure

Responsibilities

  • Ensure uptime, performance, and capacity planning for a large scale big data production platform
  • Leverage monitoring tools to proactively detect and resolve issues
  • Lead triage, troubleshooting, root cause analysis, and post-incident reviews
  • Define and track reliability metrics (SLIs & SLOs)
  • Participate in release planning, scrums, and cross-team coordination in a SAFe Agile environment

Benefits

  • 3 Medical Plan Options
  • Dental and Vision coverage
  • Flexible Savings Account (FSA) and Health Savings Account (HSA)
  • Short-Term & Long-Term Disability insurance
  • Employee Assistance Program (EAP)
  • Training and Educational Assistance
  • Paid Time Off (PTO)
  • 11 Federal holidays
  • 401k plan with up to a 6% match and immediate vesting
Full Job Description
Location: Bethesda, MD

Category: Software Engineer
Travel Required: No
Remote Type: Hybrid
Clearance: TS/SCI

As a Platform Operations Engineer you will work with a team to ensure the availability, reliability, and performance of a full stack, containerized microservices platform. You also will partner with a multidisciplinary team of systems engineers, developers, integrators, and system administrators in the following areas:
  • System Reliability & Performance - Ensuring uptime, performance, and capacity planning for a large scale big data production platform with a microservice architecture running on Kubernetes, Elasticsearch, PostgreSQL, Kafka, and technologies such as Java, Python, React, and low code tools like Appian
  • Monitoring & Observability - Leveraging monitoring tools to proactively detect and resolve issues
  • Incident Response - Leading triage, troubleshooting, root cause analysis, and post incident reviews
  • SLIs & SLOs - Defining and tracking reliability metrics
  • SAFe Agile - Participating in release planning, scrums, design sessions, bug triage, and cross team coordination


You bring enthusiasm, the ability to work well with people from different disciplines with varying degrees of technical experience, and meet the following qualifications:
  • BS in Engineering, Computer Science, Systems Engineering, or related field (or equivalent experience) with 8+ years of relevant experience; 6+ years with a Master's; additional experience may substitute for a degree
  • Active TS/SCI clearance with the ability to obtain and maintain a polygraph
  • At least one DoD 8570.01 M IAT Level II+ certification (e.g., Security+ CE, CySA+, CCNA Security, SSCP, CISSP (or Associate))
  • Ability to obtain Privileged User Account (PUA) certification
  • Experience with Kubernetes, GitLab pipelines, Linux, and containerized environments
  • Experience supporting enterprise scale production systems
  • Experience with cloud services (preferably AWS) and cloud infrastructure
  • Familiarity with Elasticsearch, PostgreSQL, Logstash, Kibana, and Keycloak
  • Demonstrated success in cross functional coordination and execution
  • Strong communication skills and the ability to perform under pressure during incidents


You will stand out even more if you bring:
  • Experience with Agile methodologies
  • Experience with creating customized dashboards to track SLIs and other key performance indicators
  • Development experience (Bash, PowerShell, SALT, Python, Groovy, Java, etc.)
  • Experience with Appian or other low-code platforms
  • Experience with technologies such as Kafka, AMQP/JMS, Prometheus/Grafana, GPU-based Kubernetes, SALT automation, Nexus, or GraphQL
  • Knowledge of security best practices (authN/Z, secrets management, data protection)
  • Infrastructure-as-code experience (CloudFormation, Terraform, Pulumi)
  • AWS cloud certifications


Pay Rate

Salary range considers factors such as (but not limited to) scope and responsibilities of the position, candidate's work experience, education/ training, key skills, as well as market and business considerations when extending an offer.

Benefits
  • 3 Medical Plan Options
  • Dental and Vision
  • FSA, DCFSA, HSA
  • Life/AD&D Insurance
  • Short-Term & Long-Term Disability
  • Employee Assistance Program (EAP)
  • Training and Educational Assistance
  • Paid Time Off (PTO)
  • 11 Federal holidays
  • 401k plan with up to a 6% match (100% immediate vesting)


Similar Jobs

More Jobs at Sunayu, LLC

More Aerospace & Defense Jobs

Find similar Platform Operations Engineer - TS/SCI jobs: