Senior Production Engineer, Managed Cloud

Crusoe

$170K — $205K *
Enterprise Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience in distributed systems and cloud services
  • Proven software engineering skills beyond scripting
  • Experience in defining and measuring Service Level Indicators (SLIs) and Service Level Objectives (SLOs)
  • Familiarity with monitoring and observability practices
  • Proficiency in one or more modern programming languages such as Python, Go, Java, or C++
  • Knowledge of Kubernetes or container orchestration platforms
  • Strong communication and teamwork skills

Responsibilities

  • Design and manage AI services that ensure reliability, especially for LLM workloads
  • Set, monitor, and enhance performance and reliability metrics
  • Work alongside AI and infrastructure teams on large-scale computing clusters
  • Automate systems observability with effective telemetry solutions
  • Troubleshoot and resolve issues in distributed AI infrastructures
  • Help design the architecture for new AI-centric distributed systems

Benefits

  • Industry competitive pay
  • Restricted Stock Units in a growing tech company
  • Comprehensive health insurance options including dental and vision
  • Employer contributions to Health Savings Accounts (HSA)
  • Generous paid parental leave policy
  • Life, short-term, and long-term disability insurance
  • Access to telehealth services via Teladoc
  • 401(k) with 100% match up to 4%
  • Significant paid time off and holiday schedule
  • Cell phone reimbursement
  • Tuition reimbursement
  • Access to the Calm app for mental wellness
  • MetLife Legal services
  • Company-paid commuter benefit of $300 per month
Full Job Description
About the Role

At Crusoe, our Production Engineering team ensures the reliability and scalability of Crusoe's AI-optimized cloud platform. We're looking for a Senior Production Engineer with a strong background in distributed systems and cloud services to help us build and deliver a seamless Cloud experience at scale. This role is central to delivering highly available, performant, and cost-efficient AI infrastructure that powers compute-intensive, latency-sensitive workloads for our customers.

What You'll Be Working On
  • Design and operate reliable managed AI services with a focus on serving and scaling LLM workloads
  • Define, measure, and improve SLIs/SLOs across to ensure performance and reliability targets are met
  • Collaborate with AI, platform, and infrastructure teams to optimize large-scale training and inference clusters
  • Automate observability by building telemetry and performance tuning strategies for latency-sensitive services
  • Investigate and resolve reliability issues in distributed AI systems using telemetry, logs, and profiling
  • Contribute to the architecture of next-generation distributed systems purpose-built for AI-first environments


What You'll Bring to the Team
  • Strong software engineering background - experience building production-grade systems beyond scripting or Bash
  • Demonstrated experience in distributed systems design and implementation
  • SRE mindset and experience (whether or not under the SRE title) including:
    • Defining and measuring SLIs/SLOs
    • Building monitoring and observability systems
    • Driving performance and reliability improvements
    • Designing fault-tolerant systems and automated testing strategies
  • Proficiency in at least one modern programming language (Python, Go, Java, C++)
  • Familiarity with Kubernetes or container orchestration platforms
  • Strong collaboration and communication skills
  • Ability to thrive in a fast-paced, mission-driven environment


Benefits
  • Industry competitive pay
  • Restricted Stock Units in a fast growing, well-funded technology company
  • Health insurance package options that include HDHP and PPO, vision, and dental for you and your dependents
  • Employer contributions to HSA accounts
  • Paid Parental Leave
  • Paid life insurance, short-term and long-term disability
  • Teladoc
  • 401(k) with a 100% match up to 4% of salary
  • Generous paid time off and holiday schedule
  • Cell phone reimbursement
  • Tuition reimbursement
  • Subscription to the Calm app
  • MetLife Legal
  • Company paid commuter benefit; $300 per month


Compensation

Compensation will be paid in the range of $170,000K to $205,000. Restricted Stock Units are included in all offers. Compensation to be determined by the applicant's education, experience, knowledge, skills, and abilities, as well as internal equity and alignment with market data.

Similar Jobs

More Jobs at Crusoe

More Enterprise Technology Jobs

Find similar Senior Production Engineer, Managed Cloud jobs: