Core Scientific

Staff Site Reliability Engineer

Core Scientific$125K — $150K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science or related field, with 7+ years of experience in SRE, DevOps, or Infrastructure Engineering.
  • Broad technical experience across infrastructure and distributed systems, anticipating scalability and reliability challenges.
  • Strong understanding of distributed systems, including service-to-service communication and failure modes.
  • Experience in regulated environments, adhering to compliance and change control standards.
  • Proficient in hybrid cloud environments, with AWS preferred and necessary on-premises experience.
  • Strong expertise in Infrastructure as Code and tooling, including Terraform and Ansible.
  • Knowledge of Kubernetes and observability platforms like Datadog.

Responsibilities

  • Lead the delivery of complex technical projects from design to implementation and operation.
  • Own system design, implementation, and reliability across hybrid cloud and on-premises.
  • Be accountable for system outcomes, including reliability and performance in regulated environments.
  • Coordinate work across teams and engineers, delegating while remaining hands-on when needed.
  • Collaborate with application architecture teams to influence system design decisions.
  • Build and operate infrastructure applications using automation and Infrastructure as Code practices.
  • Enhance observability and incident response practices across systems.

Benefits

  • Mentorship opportunities to elevate technical skills and expertise across the organization.
  • Foster a culture of open and respectful communication within teams and broader organization.
  • Exposure to diverse technical challenges in a hybrid work environment.
  • Possibility for occasional travel to data centers, providing exposure to real-world operational conditions.
  • Structured work schedule with a clear on-call rotation and defined working hours.
Full Job Description
The Job
We are seeking a capable, motivated generalist who thrives in a change-controlled, compliant environment and enjoys working across hybrid cloud and on-premises systems. This role partners closely with application architecture and peer engineering teams while contributing hands-on across platform engineering, DevOps, and SRE.

This position is expected to take ownership of complex technical initiatives and see them through to completion-balancing hands-on implementation with effective delegation and cross-team coordination.
Responsibilities
  • Lead end-to-end delivery of complex technical initiatives, from problem definition and design through implementation, rollout, and operation.
  • Own the design, implementation, and reliability of systems across hybrid cloud and on-premises environments.
  • Take accountability for technical outcomes, including system reliability, scalability, and performance in regulated, change-controlled environments.
  • Drive execution by coordinating work across engineers and teams, delegating effectively while remaining hands-on where needed.
  • Partner with application architecture and peer teams to shape system design and influence technical decisions.
  • Build, deploy, and operate infrastructure and applications using automation and infrastructure as code.
  • Implement secure, immutable infrastructure using modern tooling (e.g., Terraform, Kubernetes, Helm, Ansible).
  • Improve observability, monitoring, and incident response practices.
  • Establish and promote best practices for reliability, security, and operational excellence across teams.
  • Mentor engineers and contribute to raising the technical bar across the organization.
  • Foster open, respectful, and professional communication directly within the team as well as with co-workers/ teammates and leaders across the organization.
  • Performs other duties as assigned.


Qualifications
  • Bachelor's degree in Computer Science or a related field, 7+ years of experience, or equivalent demonstrated impact in SRE, DevOps, or Infrastructure Engineering.
  • Broad technical experience across infrastructure and distributed systems, with the ability to design effective solutions, apply appropriate patterns, and anticipate scaling, reliability, and operational challenges.
  • Strong understanding of distributed systems behavior, including application runtime characteristics, service-to-service communication, networking, and failure modes in production environments.
  • Experience operating in regulated, compliant, or change-controlled environments.
  • Experience working in hybrid environments (AWS preferred; on-premises infrastructure required).
  • Strong experience with Infrastructure as Code, configuration management, and orchestration tools (Terraform, Helm, Kustomize, Ansible).
  • Experience with Kubernetes and virtualization technologies.
  • Experience with observability platforms (e.g., Datadog), including building monitoring and alerting integrations.
  • Experience with build and release systems (e.g., GitHub Actions, Makefiles, Python tooling).


Reports To
Site Reliability Engineering Manager

Location
To be considered for the role you must reside near Miami, FL or Austin, TX.

Travel
Occasional travel may be required as needed.

Work Environment
This job typically operates in a professional office environment and routinely utilizes standard equipment, including laptop computers and smartphones. This role may also travel to data center sites, and the work environment at a data center may contain loud noise, construction, and other operational elements.

Physical Demands
While performing the duties of this job, the employee is frequently required to sit, stand, walk, use hands, and lift up to 25 pounds.

Position Type/ Expected Hours of Work
This is a full-time position. General hours and days of work are Monday through Friday, 8:00 a.m. to 5:00 p.m. The employee is expected to be available generally around U.S. time zones and will be part of an on-call rotation. The current rotation is 1 week every 5 weeks.

Supervisory Experience (Yes or No)
No

About Core Scientific

Core Scientific is a blockchain and artificial intelligence infrastructure and software solutions provider. The company offers a range of products and services, including blockchain hosting and colocation, artificial intelligence and machine learning, and software development. Core Scientific's clients include some of the world's largest blockchain and artificial intelligence companies, as well as Fortune 500 companies. The company was founded in 2017 and is headquartered in Seattle, Washington.
Learn more about Core Scientific
Size
500 employees
Market Cap
$26 million
Industry
NASDAQ

Similar Jobs

More Information Technology Jobs

Find similar Staff Site Reliability Engineer jobs: