Software Engineer, Infrastructure & Reliability

CrewAI

$120K — $145K *
US-AnywhereRemote in United States
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of infrastructure/platform engineering experience in production SaaS environments
  • Proficient with AWS, Docker, CI/CD, GitHub Actions, and container services
  • Experience with ECS and/or Kubernetes; Helm knowledge is advantageous
  • Proficient in PostgreSQL, Redis, job systems, queues, and web services in production
  • Strong debugging skills across application, infrastructure, network, deployment, and dependencies
  • Security-focused in IAM, secrets, workload identity, and production access
  • Capable of writing reliable automation using Python, Ruby, Go, Bash, or similar languages

Responsibilities

  • Own and enhance CrewAI's platform infrastructure across AWS, ECS/ECR, Docker, Kubernetes, networking, and databases
  • Construct and maintain CI/CD pipelines for various deployment activities and ensure safe rollbacks
  • Enhance reliability in cloud and enterprise deployments through incident response and operational practices
  • Collaborate with runtime engineers on workload optimizations and partner with product teams on deployment behaviors
  • Manage observability and telemetry for production systems, implementing logs, metrics, and alerting
  • Fortify security and compliance for IAM, secrets management, and vulnerability scanning
  • Develop automation tools for field engineers and customer self-hosting initiatives
  • Minimize operational toil by automating repetitive workflows and improving deployment processes

Benefits

  • Flexible work environment with remote options
  • Opportunities for professional development and growth
  • Engagement with cutting-edge technology and practices in cloud infrastructure
  • Exposure to advanced security and compliance challenges
  • Collaborative culture emphasizing teamwork and shared success
Full Job Description
The Role

You'll build and operate the platform infrastructure behind CrewAI's cloud and enterprise deployments. You'll work across multiple hyperscalers - AWS, Azure, and GCP. You'll work on containers, CI/CD, deployment automation, observability, secrets, networking, and runtime reliability. Your job is to make the product and runtime teams faster while making customer's production environments safer.

This is not a pure DevOps support role. You'll write code, improve systems, design deployment paths, harden production, and build the internal platform that lets CrewAI scale and scale our customer deployments.
What You'll Do
  • Own and improve the infrastructure that runs CrewAI's platform: AWS, ECS/ECR, Docker, Kubernetes/Helm, networking, secrets, databases, Redis, and related services.
  • Build and maintain CI/CD pipelines for build, test, image publishing, migrations, environment promotion, rollbacks, and deploy safety.
  • Improve reliability across cloud and enterprise deployments: health checks, alerting, incident response, capacity planning, recovery paths, and operational runbooks - and own the front-line on-call rotation and its SLAs.
  • Partner with runtime engineers on Celery/FastAPI/Redis workloads and with product engineers on Rails/Solid Queue/Postgres production behavior.
  • Manage production observability and telemetry infrastructure: logs, metrics, traces, dashboards, Sentry/OpenTelemetry plumbing, actionable alerts, and telemetry export to customers' own monitoring systems.
  • Harden security and compliance posture across IAM, workload identity, secrets management, vulnerability scanning, dependency/image hygiene, and least-privilege access.
  • Build the tooling and automation that lets field engineers and customers run self-hosted installs themselves - Helm charts, environment config, release artifacts, pre-flight checks, and install runbooks - so engineering does fewer hands-on installs over time.
  • Reduce operational toil by automating recurring workflows and making deployments boring.

Requirements
What We're Looking For
  • Strong infrastructure/platform engineering experience in production SaaS environments.
  • Deep practical experience with AWS, Docker, CI/CD, GitHub Actions, and containerized services.
  • Experience with ECS and/or Kubernetes; Helm experience is a strong plus.
  • Comfort operating PostgreSQL, Redis, background job systems, queues, and web services in production.
  • Strong debugging instincts across app, infra, network, deploy, and dependency layers.
  • Security-minded approach to IAM, secrets, workload identity, vulnerability management, and production access.
  • Ability to write reliable automation in Python, Ruby, Go, Bash, or similar.
  • Calm, rigorous approach to incidents, rollbacks, migrations, and production change management.
Bonus
  • Experience with AI/agent platforms, workflow runtimes, or high-volume async execution systems.
  • Experience supporting enterprise/self-hosted deployments.
  • Terraform or other IaC experience.
  • SRE background: SLOs, incident review, capacity planning, load testing.
  • Familiarity with Rails, FastAPI, Celery, OpenTelemetry, or multi-service observability.

Similar Jobs

More Jobs at CrewAI

More Information Technology Jobs

Find similar Software Engineer, Infrastructure & Reliability jobs: