PathAI

Senior/Staff Site Reliability Engineer - Data Center

PathAI$146K — $225K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • BS in Computer Science, Engineering, or a related field.
  • 8 years of experience in facilities, networking, automation, or related roles.
  • Experience with modern datacenter network designs across various network layers.
  • Hands-on administration of physical hardware in production environments.
  • Familiarity with virtualization, containerization, or orchestration platforms.
  • Strong expertise in high-performance storage solutions.
  • Proficiency with automation tools like Ansible or RedFish.

Responsibilities

  • Implement SRE best practices to enhance operational efficiency.
  • Design and operate data centers for the Machine Learning team.
  • Build secure on-premises environments complying with NIST/ISO standards.
  • Integrate on-premises and cloud infrastructure for a hybrid cloud setup.
  • Conduct root-cause analysis to improve infrastructure reliability.
  • Participate in on-call rotations and respond to urgent incidents.

Benefits

  • Comprehensive health benefits package.
  • Flexible hybrid work arrangement with remote options for exceptional candidates.
  • Opportunities for professional development and continuous learning.
Full Job Description
We are expanding our team and recruiting for a skilled Senior/Staff Site Reliability Engineer focused on designing, building, and operating our on-prem/cloud environment.
The Opportunity
  • You will advance the state of our operations by implementing SRE best practices - focusing on users, monitoring, and automation.
  • You will design, build and operate our data center to support our rapidly growing Machine Learning team.
  • You will build highly-secure on-premises environments handling NIST/ISO standards.
  • You will integrate on-premises datacenter environments with existing cloud infrastructure to create a seamless hybrid cloud environment.
  • You will improve the reliability and resilience of our infrastructure through root-cause analysis and reviewing gaps in designs, and implementations of our infrastructure.
  • You will participate in platform on-call rotations and assist with urgent incident response.
Who You Are:(Required)
  • You have a BS in Computer Science, Computer Engineering, Electrical Engineering, Software Engineering or closely related technical field.
  • You have 8 years experience working in physical hardware/facilities, networking, automation or other relevant areas.
  • You have demonstrated experience with modern datacenter network designs and comfort operating across network layers.
  • You've administered physical hardware stacks in production settings (iDRAC/IPMI/Nvidia UFM/Juniper Systems).
  • You have demonstrated experience and opinions on virtualization, containerization, or container orchestration platforms. (EKS-Anywhere/ClusterAPI/KVM).
  • You have strong expertise in storage solutions and optimizing them for high-performance workloads (e.g., Quobyte, S3, FSx, EFS).
  • You are highly proficient with automation tools; you eliminate toil by automating everything through scripting, configuration management tools (Ansible/RedFish).
  • You've built monitoring infrastructure with modern observability tools (Datadog/Grafana/Prometheus).
  • You have a proven operational background managing critical production systems, with extensive experience in incident response, infrastructure scaling, and navigating high-growth challenges.
  • You have the ability to travel to onsite Datacenter location(s) as needed.

Preferred:
  • You have outstanding interpersonal, verbal, and written communication and influencing skills: have built and cultivated important relationships both inside and outside of the organization and externally; have proven abilities to influence internal partners and stakeholders, thought leaders, national advocacy organizations, national standard-setting bodies, and other relevant external parties.
  • You have strong analytical and critical thinking skills with attention to detail; you have the ability to manage multiple projects and drive results in a fast-paced environment; you have a collaborative mindset with demonstrated leadership capabilities..

This is a hybrid position based in Boston (preferred), NYC, Indianapolis. (A remote option may be considered for an exceptional candidate.)
Relocation benefits are not available for this position.

The expected salary range for this position based on the primary location Boston, MA is $146,250 - $225,000. Actual pay will be determined based on experience, qualifications, geographic location, and other job-related factors permitted by law.

About PathAI

PathAI is a healthcare technology company that uses artificial intelligence and machine learning to improve the accuracy and speed of pathology diagnoses. The company's platform analyzes digital pathology images to help pathologists make more accurate diagnoses and improve patient outcomes. PathAI's technology has applications in cancer diagnosis and treatment, drug development, and clinical trials. The company aims to improve the quality of healthcare by providing more accurate and efficient pathology services.
Learn more about PathAI
Size
100 employees
Industry
Founded
2016

Similar Jobs

More Jobs at PathAI

More Information Technology Jobs

Find similar Senior/Staff Site Reliability Engineer - Data Center jobs: