OverviewREI Systems is seeking a DevOps Engineer to support the design, implementation, automation, and operation of enterprise-scale cloud environments and software delivery platforms. This individual will work closely with software engineering, infrastructure, security, and operations teams to build reliable CI/CD pipelines, automate cloud infrastructure, support containerized applications, and improve the performance and availability of production and non-production environments.
The ideal candidate brings hands-on experience with AWS, Kubernetes, CI/CD, Infrastructure as Code, scripting, monitoring, and production support within complex enterprise environments.
DevOps Engineering & Automation
- Design, develop, maintain, and improve CI/CD pipelines and software delivery tools supporting enterprise-scale applications and development teams.
- Design, develop, and maintain Infrastructure as Code (IaC) solutions for cloud infrastructure and application environments.
- Build, deploy, and support containerized applications and microservices using Kubernetes, Docker, AWS EKS, and/or OpenShift.
- Develop automation scripts using Bash, Python, Groovy, or similar scripting languages to streamline DevOps and operational processes.
- Support application and microservice deployments across cloud-based development, test, staging, and production environments.
- Partner with development teams to troubleshoot application, configuration, deployment, and infrastructure-related issues.
- Contribute to short- and long-term strategies for improving cloud scalability, resiliency, automation, and operational efficiency.
- Evaluate and implement emerging DevOps, cloud-native, and automation technologies to continuously improve engineering practices.
Production Operations & Reliability
- Support the availability, reliability, and performance of 24x7 production and non-production environments.
- Participate in incident, problem, and change management processes, including root cause analysis and corrective action planning.
- Monitor infrastructure, applications, networks, servers, and databases to proactively identify performance or availability issues.
- Develop and track operational KPIs related to system availability, uptime, SLA performance, incidents, and operational health.
- Perform operational maintenance activities including operating system patching, security remediation, vulnerability management, and environment upgrades.
- Administer and improve monitoring, alerting, logging, and observability capabilities.
- Analyze logs, metrics, and monitoring data to identify root causes of critical incidents and recurring system issues.
- Participate in an on-call rotation and provide hands-on support during production incidents, outages, deployments, and service transitions.
- Continuously improve system observability through enhanced monitoring, dashboards, metrics, alerts, and log analytics.
Qualifications
- 3+ years of relevant DevOps, Cloud Engineering, Site Reliability Engineering, or Systems Engineering experience.
- Strong understanding of CI/CD concepts and pipelines, preferably using Jenkins and Groovy.
- Hands-on experience with container technologies including Kubernetes, Docker, AWS EKS, and/or OpenShift.
- Strong understanding of microservices architecture and container-based application environments.
- Experience developing automation scripts using Bash, Python, Perl, or similar languages.
- Hands-on experience with AWS services including VPC, EC2, EKS, RDS, IAM, CloudWatch, and related cloud services.
- Experience with monitoring, logging, and application performance tools such as AWS CloudWatch, New Relic, or similar platforms.
- Experience with source code and version control systems such as Git and/or Subversion.
- Experience supporting or maintaining Java and/or Node.js web applications, including technologies such as Spring Boot and React.
- Experience troubleshooting production application, infrastructure, networking, and deployment issues.
- Understanding of incident management, problem management, change management, SLAs, and production support practices.
- Strong analytical, troubleshooting, communication, and teamwork skills.
Education: Bachelor’s degree in computer science or a related field.
Clearance: Candidate must be a US Citizen to support this federal project and able to obtain and maintain a Clearance.
Location: Hybrid (2 Day per week in our Sterling, VA HQ)
#LI-HYBRID
#LI-TK1