Position Description:We are seeking a Senior DevOps Engineer to design, automate, and support our AWS cloud infrastructure and Kubernetes-based application platform. This role will focus on reliability, scalability, security, observability, and deployment automation across production environments.
Position Location:Remote - New York, Texas, Philadelphia, Florida, North Carolina, Minnesota, Colorado, Georgia, Illinois
Reports To:Sr. Manager, DevOps and Infrastructure
What We're Looking For:• 5+ years of experience in DevOps, Infrastructure Engineering, Site Reliability Engineering, Cloud Engineering, or a similar role.
• Strong hands-on experience with AWS cloud infrastructure.
• Production experience with Kubernetes, preferably Amazon EKS.
• Experience with Infrastructure as Code using OpenTofu, Terraform, or similar tools.
• Experience with Argo CD, Helm, and GitOps-based deployment workflows.
• Strong scripting or development experience with Python.
• Experience supporting Aurora PostgreSQL, ElastiCache Redis, and Amazon MQ for RabbitMQ in production environments.
• Experience with Datadog or similar observability platforms.
• Strong understanding of cloud networking, including VPCs, subnets, routing, security groups, DNS, TLS, ingress, and load balancing.
• Comfortable participating in an on-call rotation and supporting production incident response.
• Strong troubleshooting, communication, and collaboration skills.
Additional skills:
• AWS, Amazon EKS, Aurora PostgreSQL, ElastiCache Redis, Amazon MQ for RabbitMQ, AWS Load Balancers, IAM, VPC networking.
• OpenTofu, Terraform, Infrastructure as Code, Kubernetes, Helm, containers, ingress, services, ConfigMaps, Secrets, autoscaling.
• Argo CD, GitOps, CI/CD, release automation, Python, Bash or shell scripting, Datadog, metrics, logs, tracing, dashboards, alerting.
• Experience supporting high-availability SaaS or customer-facing platforms.
• Experience with Kubernetes autoscaling, ingress controllers, cluster upgrades, and resource optimization.
• Experience with backup, restore, disaster recovery, and capacity planning.
• Familiarity with AWS Well-Architected principles and cloud security best practices.
• Experience with GitHub Actions, GitLab CI, Jenkins, or similar CI/CD tools.
• Experience supporting message broker platforms such as RabbitMQ or Amazon MQ.
• AWS or Kubernetes certifications are a plus.
• Strong ownership mindset, operational discipline, and a passion for automation, reliability, and continuous improvement.
• Ability to communicate clearly with technical and non-technical stakeholders, with a willingness to mentor others.
Unleash your potential: What you will be doing and owning:• Build and maintain AWS infrastructure using OpenTofu.
• Operate Kubernetes workloads on Amazon EKS.
• Manage GitOps deployments using Argo CD and Helm.
• Develop automation and operational tooling using Python and scripting languages.
• Support AWS services including Amazon EKS, Aurora PostgreSQL, ElastiCache Redis, Amazon MQ for RabbitMQ, load balancers, IAM, VPC networking, DNS, and security groups.
• Implement monitoring, dashboards, alerts, logs, metrics, and tracing using Datadog.
• Troubleshoot production issues and support incident response and root cause analysis.
• Participate in an on-call rotation to support production systems, respond to incidents, and assist with after-hours maintenance or escalations as needed.
• Improve CI/CD pipelines, release automation, and deployment reliability.
• Partner with engineering and security teams to improve infrastructure standards, operational readiness, and cloud security.
• Create and maintain technical documentation, runbooks, and operational procedures.
Interview Process:- Interview #1: Video Screen with Talent Acquisition Team
- Interview #2: Video interview with the Hiring Manager (via MS Teams)
- Interview #3: Video interview with the Team (via MS Teams)