Senior DevOps Engineer

Akkadian Labs

$120K — $145K *
Technical Services
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 10+ years of experience in DevOps or Site Reliability Engineering (SRE)
  • Expertise with AWS services including EC2, ECS, S3, IAM, Lambda, and CloudWatch
  • Proficient in infrastructure-as-code tools like Terraform and CloudFormation
  • Strong understanding of Linux environments
  • Experienced in containerization with Docker and Kubernetes
  • Scripting skills in Python, Bash, or similar languages
  • Proven experience with CI/CD pipelines and related tools
  • Familiar with observability tools like Prometheus, Grafana, and ELK
  • Experience implementing secure DevOps practices and compliance standards
  • Background in supporting AI or machine learning workloads

Responsibilities

  • Deploy and maintain scalable infrastructure in AWS and hybrid cloud environments
  • Manage infrastructure-as-code using Terraform or CloudFormation
  • Maintain Linux-based environments
  • Design and implement containerization using Docker and orchestration with Kubernetes
  • Design and manage AI agent workloads, including provisioning compute resources
  • Build and maintain model deployment pipelines for AI models
  • Manage monitoring and observability using tools like Prometheus and Grafana
  • Troubleshoot system issues and contribute to incident response
  • Build, maintain, and optimize CI/CD pipelines
  • Automate routine operational tasks and collaborate with engineering teams
  • Implement security controls and support compliance initiatives
  • Document infrastructure, processes, and operational procedures

Benefits

  • Fully remote work environment
  • Competitive benefits package including medical, dental, and vision coverage
  • Company-paid life insurance and disability policies
  • 401(k) with generous matching program
  • Paid time off
Full Job Description
Who You Are

You are a hands-on Senior DevOps Engineer with deep experience designing, building, and strengthening the infrastructure that keeps platforms secure, compliant, and uncomplicated. You bring substantial experience maintaining production infrastructure and improving operational reliability through disciplined engineering practice. You have a track record of helping to scale systems with a focus on dependability, observability, deployment confidence, and repeatable processes that support long-term growth. You are energized by solving infrastructure challenges in complex enterprise environments and by building systems that reduce manual work, improve governance, and support operational accuracy at scale.

What You'll Do

This role supports Akkadian's purpose to uncomplicate work by removing friction between people and their purpose - helping enterprise customers govern, automate, and continuously control complex collaboration environments across cloud, hybrid, and on-premises systems. You will design, implement, and maintain scalable and secure infrastructure and DevOps processes at Akkadian Labs. You will work closely with development, QA, and product teams to enable reliable deployments, automate workflows, and improve system observability across Rocky OS-based, AWS-hosted, and on-premises solutions.

This is a hands-on technical role focused on system-level execution, continuous improvement, and operational excellence within the DevOps function.

Key Responsibilities

Infrastructure and Environment Management
  • Deploy and maintain scalable infrastructure in AWS and hybrid cloud environments.
  • Manage infrastructure-as-code (IaC) using Terraform, CloudFormation, or similar tools.
  • Maintain Linux-based environments.
  • Design and implement containerization using Docker and orchestration via Kubernetes.

AI and Agent Infrastructure Implementation & Support
  • Design, deploy and manage AI agent workloads, including provisioning compute instances and managing resource scaling for inference-heavy tasks.
  • Build and maintain model deployment pipelines, including versioning, testing, and rollback of AI models in production environments.
  • Monitor AI API consumption and infrastructure costs, implementing alerting and controls to prevent runaway usage and support budget visibility.
  • Collaborate with the engineering team to implement infrastructure-level security guardrails for AI systems, including access controls and data isolation for model inputs and outputs.

Observability and Reliability
  • Manage monitoring and observability efforts using tools such as Prometheus, Grafana, and the ELK stack.
  • Troubleshoot system issues and contribute to incident response and root cause analysis.
  • Develop and execute strategies for improving system reliability, performance, and uptime.

CI/CD and Automation
  • Build, maintain, and optimize CI/CD pipelines using tools such as Jenkins, Bitbucket CI/CD, or similar.
  • Automate routine operational tasks including builds, testing, deployments, and system updates.
  • Collaborate with engineering teams to integrate pipelines with Akkadian tools.

Security and Compliance
  • Follow secure DevOps practices and implement and maintain security controls.
  • Support compliance initiatives and vulnerability remediation efforts.

Collaboration and Documentation
  • Work closely with DevOps, engineering, QA, and product teams to support deployments and releases.
  • Maintain documentation for infrastructure, processes, and operational procedures.
  • Participate in collaborative team processes and continuous improvement initiatives.


Requirements
  • Experience: 10+ years of experience in DevOps or Site Reliability Engineering (SRE).
  • Cloud Expertise: Expertise with AWS (e.g., EC2, ECS, S3, IAM, Lambda, CloudWatch).
  • Infrastructure as Code: Expertise with infrastructure-as-code tools, including Terraform and CloudFormation.
  • Linux Knowledge: Strong knowledge of Linux environments.
  • Containerization: Experienced with Docker and Kubernetes.
  • Scripting: Scripting ability in Python, Bash, or similar languages.
  • CI/CD: Experience building or maintaining CI/CD pipelines and related tools.
  • Observability: Experience in monitoring and observability tools such as Prometheus, Grafana, and ELK.
  • Security: Experience in implementing secure DevOps practices and compliance frameworks (SOC2, ISO, etc).
  • Experience supporting AI or machine learning workloads and compute environments.
  • Exposure to AI model deployment pipelines and model versioning practices.
  • Familiarity with hybrid cloud or on-premises environments.
  • Exposure to security best practices in DevOps contexts, including AI-specific concerns such as data isolation and access controls.
  • Experience supporting production systems and participating in on-call rotations.

Benefits

We offer a fully remote environment, plus a competitive benefits package including medical, dental, vision, company-paid life insurance and disability policies, 401(k) with a generous matching program, and paid time off.

Similar Jobs

More Jobs at Akkadian Labs

  • Senior DevOps Engineer
    $120K — $145K *
    Hoboken, NJ 07030 (Hudson County)
    Technical Services
    In-Person
  • Account Manager
    $80K — $95K *
    Hoboken, NJ 07030 (Hudson County)
    Enterprise Technology
    In-Person

More Technical Services Jobs

Find similar Senior DevOps Engineer jobs: