Job Summary- Seeking a Site Reliability Engineer to support a cloud-based platform built with Java and Free and Open-Source Software technologies, including Kubernetes, Hadoop, and Accumulo
- The platform enables the execution of data-intensive analytics across managed infrastructure in a mission-focused environment
- The ideal candidate is self-motivated, detail-oriented, and able to thrive in a fast-paced team environment while proactively identifying and resolving operational issues
- This is an on-call position and includes Tier 1 through Tier 3 support responsibilities
Primary Responsibilities- Support the operation, administration, reliability, and availability of cloud-based infrastructure and platform services
- Provide Tier 1 through Tier 3 technical support for operational issues affecting users, applications, and infrastructure
- Troubleshoot and resolve complex system, application, networking, and infrastructure issues within Linux environments
- Monitor system health, availability, performance, and operational status
- Support distributed computing technologies including Kubernetes, Hadoop, and Accumulo
- Support containerized applications and services using Docker and Kubernetes
- Develop and maintain automation and administrative scripts using Python, Bash, or similar scripting languages
- Support Hadoop Distributed File System (HDFS) environments
- Use monitoring and observability tools such as Prometheus and Grafana to identify and resolve operational issues
- Support configuration management and automation tools such as Salt and Ansible
- Participate in incident response, root cause analysis, corrective actions, and continuous improvement activities
- Support virtualization, cloud, and hybrid infrastructure environments
- Document troubleshooting procedures, operational processes, system changes, and recurring issues
- Collaborate with developers, system administrators, engineers, and mission stakeholders to maintain reliable and scalable platform operations
- Participate in on-call support and respond to operational issues as required
Required Qualifications- Must have active Top Secret/SCI clearance with NSA Full Scope Polygraph
- A Bachelor's degree in Computer Science or a related technical field is highly desired and may be considered equivalent to two (2) years of experience
- A Master's degree in a technical field may be considered equivalent to four (4) years of experience
- Degrees in Mathematics, Information Systems, Engineering, or similar disciplines will be considered technical degrees
- Fourteen (14) years of relevant technical experience
- Strong experience troubleshooting operational issues in Linux environments
- DoD 8570 IAT Level I certification or higher
- Ability to provide Tier 1 through Tier 3 support in a mission-critical environment
- Candidates must possess at least one of the following certifications:
- AWS Certified Developer - Associate
- AWS Certified Solutions Architect - Associate
- AWS Certified Solutions Architect - Professional
- AWS Certified SysOps Administrator - Associate
- Certified Kubernetes Application Developer (CKAD)
- Elastic Certified Engineer
- Elastic Certified Observability Engineer
Desired Qualifications- Experience with one or more of the following technologies is beneficial:
- Docker
- Kubernetes
- Hadoop
- Apache Accumulo
- Hadoop Distributed File System (HDFS)
- Python
- Bash
- Prometheus
- Grafana
- JIRA
- Salt
- Ansible
- Virtualization technologies
- OpenStack
- Amazon Web Services (AWS)
Exempt hourly position. 11 paid holidays, minimum of 3 weeks PTO, company sponsored group medical plan, company paid dental, vision, life insurance, and STD/LTD plans. Salary is dependent upon the candidate's experience and qualifications.
The pay range for this role is:
165,000 - 230,000 USD per year (NBP)