PathAI

Senior/Staff Site Reliability Engineer - Data Center

PathAI$146K — $225K *
Enterprise Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • BS in Computer Science, Computer Engineering, Electrical Engineering, Software Engineering or a closely related field.
  • 8 years experience in physical hardware/facilities, networking, automation, or related areas.
  • Experience with modern datacenter network designs across various network layers.
  • Administered physical hardware stacks in production settings (e.g., iDRAC, IPMI, Nvidia UFM, Juniper Systems).
  • Experience with virtualization, containerization, or container orchestration platforms (EKS-Anywhere, ClusterAPI, KVM).
  • Strong expertise in optimizing storage solutions for high-performance workloads (e.g., Quobyte, S3, FSx, EFS).
  • Highly proficient with automation tools and scripting, including configuration management tools (Ansible, RedFish).

Responsibilities

  • Implement SRE best practices with a focus on users, monitoring, and automation.
  • Design, build, and operate data center infrastructure for a growing Machine Learning team.
  • Build secure on-premises environments compliant with NIST/ISO standards.
  • Integrate on-premises data centers with cloud infrastructure to create a seamless hybrid environment.
  • Enhance infrastructure reliability through root-cause analysis and design reviews.
  • Participate in on-call platform rotations for urgent incident response.

Benefits

  • Hybrid work model with locations in Boston, NYC, and Indianapolis.
  • Remote work option for exceptional candidates.
  • Opportunity to work on cutting-edge technology supporting AI initiatives.
Full Job Description
We are expanding our team and recruiting for a skilled Senior/Staff Site Reliability Engineer focused on designing, building, and operating our on-prem/cloud environment.
The Opportunity
  • You will advance the state of our operations by implementing SRE best practices - focusing on users, monitoring, and automation.
  • You will design, build and operate our data center to support our rapidly growing Machine Learning team.
  • You will build highly-secure on-premises environments handling NIST/ISO standards.
  • You will integrate on-premises datacenter environments with existing cloud infrastructure to create a seamless hybrid cloud environment.
  • You will improve the reliability and resilience of our infrastructure through root-cause analysis and reviewing gaps in designs, and implementations of our infrastructure.
  • You will participate in platform on-call rotations and assist with urgent incident response.
Who You Are:(Required)
  • You have a BS in Computer Science, Computer Engineering, Electrical Engineering, Software Engineering or closely related technical field.
  • You have 8 years experience working in physical hardware/facilities, networking, automation or other relevant areas.
  • You have demonstrated experience with modern datacenter network designs and comfort operating across network layers.
  • You've administered physical hardware stacks in production settings (iDRAC/IPMI/Nvidia UFM/Juniper Systems).
  • You have demonstrated experience and opinions on virtualization, containerization, or container orchestration platforms. (EKS-Anywhere/ClusterAPI/KVM).
  • You have strong expertise in storage solutions and optimizing them for high-performance workloads (e.g., Quobyte, S3, FSx, EFS).
  • You are highly proficient with automation tools; you eliminate toil by automating everything through scripting, configuration management tools (Ansible/RedFish).
  • You've built monitoring infrastructure with modern observability tools (Datadog/Grafana/Prometheus).
  • You have a proven operational background managing critical production systems, with extensive experience in incident response, infrastructure scaling, and navigating high-growth challenges.
  • You have the ability to travel to onsite Datacenter location(s) as needed.

Preferred:
  • You have outstanding interpersonal, verbal, and written communication and influencing skills: have built and cultivated important relationships both inside and outside of the organization and externally; have proven abilities to influence internal partners and stakeholders, thought leaders, national advocacy organizations, national standard-setting bodies, and other relevant external parties.
  • You have strong analytical and critical thinking skills with attention to detail; you have the ability to manage multiple projects and drive results in a fast-paced environment; you have a collaborative mindset with demonstrated leadership capabilities..

This is a hybrid position based in Boston (preferred), NYC, Indianapolis. (A remote option may be considered for an exceptional candidate.)
Relocation benefits are not available for this position.

The expected salary range for this position based on the primary location Boston, MA is $146,250 - $225,000. Actual pay will be determined based on experience, qualifications, geographic location, and other job-related factors permitted by law.

About PathAI

PathAI is a healthcare technology company that uses artificial intelligence and machine learning to improve the accuracy and speed of pathology diagnoses. The company's platform analyzes digital pathology images to help pathologists make more accurate diagnoses and improve patient outcomes. PathAI's technology has applications in cancer diagnosis and treatment, drug development, and clinical trials. The company aims to improve the quality of healthcare by providing more accurate and efficient pathology services.
Learn more about PathAI
Size
100 employees
Industry
Founded
2016

Similar Jobs

More Jobs at PathAI

More Enterprise Technology Jobs

Find similar Senior/Staff Site Reliability Engineer - Data Center jobs: