Senior Staff Deployment Automation Engineer

Crusoe

$250K — $300K *
Information Technology
11 - 15 years of experience
Job Overview by Ladders

Qualifications

  • 12+ years of experience in relevant fields with a Bachelor's or Master's degree in Computer Science or Electrical Engineering.
  • Proven ability to build and deploy automated integration testing for AI Cloud environments, focusing on Linux systems and distributed control planes.
  • In-depth understanding of the modern infrastructure stack, including Kubernetes, Docker, Terraform, and Postgres.
  • Expertise in CI/CD pipelines and GitLab tools for stable releases across multiple data centers.
  • Experience with configuration management systems such as Ansible, Puppet, or Chef.
  • Advanced Python and/or Bash skills for automating complex cluster-wide test scenarios.
  • Knowledge of Linux kernel internals, PCIe topology, VFIO, memory management, and GPU ecosystem familiarity.

Responsibilities

  • Own the deployment and integration testing automation for on-premise systems within the AI Cloud Stack.
  • Develop CI/CD platforms to ensure rapid testing, iteration, and deployment of low-level systems and applications.
  • Design and implement validation tests for linear scaling and stability across multi-node clusters.
  • Maintain and scale configurations using both custom and off-the-shelf tools like GitLab and Ansible.
  • Create automation for canary deployments, Blue/Green testing, and rollback procedures in production.
  • Develop automation frameworks in Python or Go for provisioning and testing multi-node environments.
  • Create automated test suites to validate CPU and GPU host performance and isolation.

Benefits

  • Competitive compensation and equity packages
  • Restricted Stock Units
  • Paid time off and leave of absence programs
  • Comprehensive health, dental & vision insurance
  • Employer contributions to HSA account
  • Paid parental leave
  • Paid life insurance and disability coverage
  • Professional development and tuition reimbursement
  • Mental health and wellness support
  • Commuter benefits
  • Cell phone stipend
  • 401(k) with company match
  • Volunteer time off
  • Global travel insurance and emergency assistance
  • Daily meals allowance
  • Additional location-specific perks and programs.
Full Job Description
About the Role:

As a Senior Staff/Principal Deployment Automation Engineer for the Compute Team, you will be responsible for deployment and testing automation of large-scale, multi-node GPU clusters. You will own the CI/CD infrastructure, including both deployment and integration testing, for a rapidly scaling fleet of virtualized GPU and CPU hosts across our AI Cloud. Your role is critical in ensuring the stability of the low-level infrastructure and enabling teams across our Cloud Infrastructure organization to quickly and reliably release, test, and deploy their artifacts across our datacenters.

San Francisco, Sunnyvale, Bellevue (Onsite)

What You'll Be Working On:
  • Deployment and Integration Testing Ownership: Completely own deployment and integration testing automation for all bare-metal, on-premise systems across Crusoe's AI Cloud Stack.
  • CI/CD Automation and Tooling: Build CI/CD platforms that enable developers to quickly test, iterate, and deploy critical, low-level systems and applications.
  • Multi-Node Scaling Validation: Design and execute large-scale validation tests across multi-node virtualized clusters to ensure linear scaling and stability of GPU workloads.
  • Configuration Management and Observability: Maintain and scale bare-metal Linux configurations using a mix of custom and off the shelf tooling such as Gitlab, Ansible, AWX, osquery, etc.
  • Deployment Orchestration: Create control applications to coordinate canary deployments on live production systems, run Blue/Green testing, and perform automatic rollback where necessary.
  • Cluster Orchestration: Develop and maintain automation frameworks in Python or Go to dynamically provision, configure, and stress-test multi-node virtualized environments.
  • Create automated test suites leveraging tools like fio, stress-ng, and iperf to ensure performance and multi-tenant isolation of CPU and GPU hosts.

What You'll Bring to the Team:
  • Education & Experience: 12+ YOE demonstrated ability to competently and independently perform responsibilities plus Bachelor's or Master's degree in Computer Science, Electrical Engineering, or a related technical field.
  • Experience building and deploying automated integration testing for an AI Cloud Environment, ranging from low-level Linux Systems up to Distributed Control Planes.
  • Working knowledge of the modern infrastructure stack, including Kubernetes, Docker, Terraform, and Postgres.
  • CI/CD & Gitlab: Intimate knowledge of CI/CD pipelines and Gitlab Tooling to enable stable infrastructure releases across multiple datacenters.
  • Configuration Management: Previous experience with at least 1-2 configuration management systems, including Ansible, Puppet, Chef, or SaltStack.
  • Automation & Scripting: Advanced proficiency in Python and/or Bash for automating complex cluster-wide test scenarios.
  • System Internals: Knowledge of Linux kernel internals, specifically PCIe topology, VFIO, and memory management (HugePages, IOMMU).
  • Distributed GPU Ecosystems: Familiarity with NVIDIA (CUDA/NCCL) and/or AMD (ROCm/RCCL) stacks in a multi-node context.
  • Networking Knowledge: Strong understanding of RDMA, RoCE, and InfiniBand protocols and their implementation in virtualized systems.

Bonus Points:
  • Experience with MNNVL (Multi-Node NVLink) or specialized AI fabric architectures.
  • Familiarity with hardware-level debugging tools and performance profilers (e.g., NVIDIA Nsight, AMD Omniperf).
  • Knowledge of containerized orchestration for GPUs (e.g., Kubernetes with specialized device plugins).


Benefits:
  • Competitive compensation and equity packages
  • Restricted Stock Units
  • Paid time off, paid holidays & leave of absence programs
  • Comprehensive health, dental & vision insurance
  • Employer contributions to HSA account
  • Paid parental leave
  • Paid life insurance, short-term and long-term disability
  • Professional development & tuition reimbursement
  • Mental health & wellness support
  • Commuter benefits (parking & transit)
  • Cell phone stipend
  • 401(k) Retirement plan with company match up to 4% of salary
  • Volunteer time off
  • Global travel insurance & emergency assistance
  • Daily meals allowance
  • Additional perks & programs specific to location


Compensation Range

Compensation will be paid in the range of up to $250,000 -$300,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicant's knowledge, education, and abilities, as well as internal equity and alignment with market data.

Similar Jobs

More Jobs at Crusoe

More Information Technology Jobs

Find similar Senior Staff Deployment Automation Engineer jobs: