OverviewSenior DevOps Engineer
REI Systems is seeking a Senior DevOps Engineer to lead the design, implementation, automation, and operation of enterprise-scale cloud environments and software delivery platforms supporting mission-critical federal applications.
This individual will serve as a senior technical contributor and work closely with software engineering, infrastructure, security, architecture, and operations teams to design and optimize CI/CD pipelines, cloud infrastructure, container platforms, Infrastructure as Code, automation, observability, and production operations.
The ideal candidate brings extensive hands-on experience with AWS, Kubernetes, CI/CD, Infrastructure as Code, scripting, monitoring, production operations, and cloud-native architectures and can independently troubleshoot complex application and infrastructure issues.
DevOps Engineering & Automation
- Design, develop, own, and continuously improve CI/CD pipelines and software delivery platforms supporting enterprise-scale applications and engineering teams.
- Architect and implement Infrastructure as Code (IaC) solutions for scalable, secure, and highly available cloud environments.
- Design, deploy, optimize, and support containerized applications and microservices using Kubernetes, Docker, AWS EKS, and/or OpenShift.
- Develop advanced automation solutions using Bash, Python, Groovy, or similar scripting languages.
- Lead application and microservice deployments across development, test, staging, and production environments.
- Partner with development teams to diagnose and resolve complex application, configuration, deployment, networking, and infrastructure issues.
- Establish and promote DevOps engineering standards, automation practices, and deployment best practices.
- Identify opportunities to improve cloud scalability, resiliency, security, automation, performance, and operational efficiency.
- Evaluate and implement emerging DevOps, cloud-native, containerization, and automation technologies.
- Provide technical guidance and mentorship to junior and mid-level DevOps engineers.
Production Operations & Reliability
- Ensure the availability, reliability, scalability, and performance of 24x7 production and non-production environments.
- Lead technical response and troubleshooting efforts during complex production incidents and outages.
- Perform and facilitate root cause analysis (RCA) and develop corrective and preventative actions.
- Design and improve monitoring, alerting, logging, and observability capabilities across cloud and application environments.
- Analyze application, infrastructure, network, and system performance to proactively identify reliability and capacity issues.
- Define and track operational KPIs related to availability, uptime, SLA performance, incidents, capacity, and system health.
- Lead or oversee operational maintenance activities including operating system patching, security remediation, vulnerability management, and environment upgrades.
- Improve system observability through dashboards, metrics, alerts, distributed logging, and application performance monitoring.
- Participate in an on-call rotation and provide senior-level technical support during critical production incidents, deployments, outages, and service transitions.
- Identify recurring operational issues and develop automation and engineering solutions that reduce manual intervention.
Qualifications
- 7+ years of relevant DevOps, Cloud Engineering, Site Reliability Engineering, Systems Engineering, or related experience.
- Advanced hands-on experience designing and maintaining enterprise CI/CD pipelines, preferably using Jenkins and Groovy.
- Strong hands-on experience with Kubernetes, Docker, AWS EKS, and/or OpenShift.
- Strong understanding of microservices architecture, distributed systems, and container-based application environments.
- Advanced experience with Infrastructure as Code, using technologies such as Terraform, CloudFormation, or similar platforms.
- Strong scripting and automation experience using Bash, Python, Groovy, or similar languages.
- Extensive hands-on experience with AWS services including VPC, EC2, EKS, RDS, IAM, CloudWatch, and related cloud technologies.
- Experience designing and supporting highly available, scalable, and secure AWS cloud environments.
- Strong experience with monitoring, logging, observability, and application performance platforms such as AWS CloudWatch, New Relic, Splunk, or similar tools.
- Strong experience with source code and version control systems such as Git.
- Experience supporting and troubleshooting Java and/or Node.js web applications, including technologies such as Spring Boot and React.
- Advanced troubleshooting skills across applications, cloud infrastructure, networking, containers, CI/CD pipelines, and production environments.
- Strong understanding of incident management, problem management, change management, SLAs, production support, and root cause analysis.
- Demonstrated ability to independently solve complex technical problems and drive issues through resolution.
- Experience mentoring engineers and contributing to technical standards and engineering best practices.
- Excellent analytical, troubleshooting, communication, and collaboration skills.
-
Education: Bachelor’s degree in computer science or a related field.
Clearance: Candidate must be a US Citizen to support this federal project and able to obtain and maintain a Clearance.
Location: Hybrid (2 Day per week in our Sterling, VA HQ)
#LI-HYBRID
#LI-KS1