5+ years of experience in deployment automation and integration testing
Bachelor's or Master's degree in Computer Science, Electrical Engineering, or related field
Proven experience in designing and deploying microservice applications in a distributed cloud environment
Strong understanding of IaaS resource management (Compute, Network, Storage)
Familiarity with Kubernetes, Docker, Terraform, and Postgres
Intimate knowledge of CI/CD pipelines and Gitlab tooling
Experience with configuration management systems like Ansible or Puppet
Knowledge of Linux kernel internals, particularly PCIe topology and memory management.
Responsibilities
Develop applications for deployment and integration testing on bare-metal systems
Build CI/CD platforms for rapid testing and deployment of low-level systems
Design and execute validation tests for multi-node GPU clusters
Maintain and scale Linux configurations using custom and off-the-shelf tools
Create control applications for canary deployments and Blue/Green testing
Develop automation frameworks in Python or Go for provisioning and stress-testing environments
Create automated test suites to ensure performance and isolation of CPU and GPU hosts.
Benefits
Competitive compensation and equity packages
Restricted Stock Units
Paid time off, holidays, and leave of absence programs
Comprehensive health, dental, and vision insurance
Employer contributions to HSA account
Paid parental leave
Paid life insurance and disability coverage
Professional development and tuition reimbursement
Mental health and wellness support
Commuter benefits
Cell phone stipend
401(k) retirement plan with company match
Volunteer time off
Global travel insurance and emergency assistance
Daily meals allowance
Location-specific perks and programs.
Full Job Description
About the Role:
As a Senior Deployment Automation Engineer for the Compute Team, you will be responsible for deployment and testing automation of large-scale, multi-node GPU clusters. You will develop the CI/CD infrastructure, including both deployment and integration testing, for a rapidly scaling fleet of virtualized GPU and CPU hosts across our AI Cloud. Your role is critical in ensuring the stability of the low-level infrastructure and enabling teams across our Cloud Infrastructure organization to quickly and reliably release, test, and deploy their artifacts across our datacenters.
What You'll Be Working On:
Deployment and Integration Testing: Develop applications and systems for deployment and integration testing on bare-metal, on-premise systems across Crusoe's AI Cloud Stack.
CI/CD Automation and Tooling: Build CI/CD platforms that enable developers to quickly test, iterate, and deploy critical, low-level systems and applications.
Multi-Node Scaling Validation: Design and execute large-scale validation tests across multi-node virtualized clusters to ensure linear scaling and stability of GPU workloads.
Configuration Management and Observability: Maintain and scale bare-metal Linux configurations using a mix of custom and off the shelf tooling such as Gitlab, Ansible, osquery, etc.
Deployment Orchestration: Create control applications to coordinate canary deployments on live production systems, run Blue/Green testing, and perform automatic rollback where necessary.
Cluster Orchestration: Develop and maintain automation frameworks in Python or Go to dynamically provision, configure, and stress-test multi-node virtualized environments.
Create automated test suites leveraging tools like fio, stress-ng, and iperf to ensure performance and multi-tenant isolation of CPU and GPU hosts.
What You'll Bring to the Team:
Education & Experience: 5+ YOE demonstrated ability to competently and independently perform responsibilities plus Bachelor's or Master's degree in Computer Science, Electrical Engineering, or a related technical field.
Experience building and deploying automated integration testing, ranging from low-level Linux Systems up to Distributed Control Planes.
Proven track record of designing and deploying microservice applications end to end in a distributed Cloud Environment.
A strong understanding of how cloud resources (Compute, Network, Storage) are abstracted and managed in an IaaS environment.
Working knowledge of the modern infrastructure stack, including Kubernetes, Docker, Terraform, and Postgres.
CI/CD & Gitlab: Intimate knowledge of CI/CD pipelines and Gitlab Tooling to enable stable infrastructure releases across multiple datacenters.
Configuration Management: Previous experience with at least 1-2 configuration management systems, including Ansible, Puppet, Chef, or SaltStack.
System Internals: Knowledge of Linux kernel internals, specifically PCIe topology, VFIO, and memory management (HugePages, IOMMU).
Bonus Points
Distributed GPU Ecosystems: Familiarity with NVIDIA (CUDA/NCCL) and/or AMD (ROCm/RCCL) stacks in a multi-node context.
Networking Knowledge: Understanding of RDMA, RoCE, and InfiniBand protocols and their implementation in virtualized systems.
Experience with MNNVL (Multi-Node NVLink) or specialized AI fabric architectures.
Familiarity with hardware-level debugging tools and performance profilers (e.g., NVIDIA Nsight, AMD Omniperf).
Knowledge of containerized orchestration for GPUs (e.g., Kubernetes with specialized device plugins).
Benefits:
Competitive compensation and equity packages
Restricted Stock Units
Paid time off, paid holidays & leave of absence programs
Comprehensive health, dental & vision insurance
Employer contributions to HSA account
Paid parental leave
Paid life insurance, short-term and long-term disability
Professional development & tuition reimbursement
Mental health & wellness support
Commuter benefits (parking & transit)
Cell phone stipend
401(k) Retirement plan with company match up to 4% of salary
Volunteer time off
Global travel insurance & emergency assistance
Daily meals allowance
Additional perks & programs specific to location
Compensation Range
Compensation will be paid in the range of up to $170,000 -$205,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicant's knowledge, education, and abilities, as well as internal equity and alignment with market data.