Site Reliability Engineer - Platform

Core Specialty

$110K — $130K *
Enterprise Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science, IT or Engineering, or equivalent experience
  • Experience with cloud platforms, preferably Microsoft Azure
  • Proficiency in Infrastructure as Code tools, particularly Terraform
  • Knowledge of CI/CD pipeline and source control systems
  • Experience with automation languages like PowerShell, Python, or Bash
  • Understanding of monitoring, observability, and troubleshooting tools
  • Strong collaboration and problem-solving skills

Responsibilities

  • Design and implement fault-tolerant, highly available cloud-native architectures
  • Define and operationalize SLOs, SLIs, and error budgets for key applications
  • Build and maintain Infrastructure as Code (IaC) using Terraform for compliant deployments
  • Develop automated remediation and self-healing capabilities to enhance system resilience
  • Establish enterprise-level monitoring frameworks with tools like Datadog and Azure Monitor
  • Drive cost optimization initiatives, focusing on resource tracking and rightsizing
  • Collaborate with application teams to integrate reliability engineering into CI/CD processes

Benefits

  • Hybrid work opportunity for improved work-life balance
  • Access to cutting-edge tools and technologies
  • Opportunity to enhance skills in cloud-native architecture and automation
  • Work within a collaborative team environment
  • Contribute to optimizations that impact organizational efficiency
Full Job Description

-

The Platform Engineer is responsible for designing, building, and maintaining enterprise technology platforms that enable teams across the organization to deliver solutions securely, efficiently, and consistently.

This role focuses on developing and supporting shared platform capabilities that power software delivery, data integration, file transfer, observability, and analytics. The Platform Engineer creates automation, self-service capabilities, reusable engineering patterns, and platform integrations that improve productivity and reduce complexity for platform consumers.

The Platform Engineering team owns the platforms themselves, ensuring they remain secure, reliable, scalable, and easy to consume. Business and operational teams use these platforms to execute their day-to-day processes.

Key Accountabilities/Deliverables:

  • Design and implement highly available, fault-tolerant architectures using cloud-native services (microservices, containers, serverless)

  • Define and operationalize SLOs, SLIs, and error budgets for critical applications and platforms

  • Build and maintain Infrastructure as Code (IaC) (Terraform) to ensure repeatable and compliant deployments

  • Develop automated remediation and self-healing capabilities to reduce MTTR and improve system resilience

  • Establish enterprise-level monitoring, logging, and observability frameworks (Datadog, Azure Monitor, CloudWatch, OpenTelemetry, Azure Application Insights)

  • Drive cost optimization (FinOps) initiatives, including resource utilization tracking and rightsizing recommendations

  • Support DR/BCP strategy execution, including failover testing and regional isolation validation

  • Collaborate with application teams to embed reliability engineering practices into CI/CD pipelines


Experience:

Applicants must be authorized to work for any employer in the U.S.  We are unable to sponsor or take over work authorization sponsorship now or in the future for this position. 

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent experience.

  • Experience working with cloud platforms, preferably Microsoft Azure.

  • Experience with Infrastructure-as-Code tools such as Terraform.

  • Experience with source control platforms and CI/CD pipelines.

  • Knowledge of automation and scripting using PowerShell, Python, Bash, or similar languages.

  • Knowledge of Power Platform 

  • Understanding of modern software delivery and DevOps practices.

  • Familiarity with monitoring, observability, and troubleshooting tools.

  • Strong problem-solving and collaboration skills.



#LI-Hybrid

-

Similar Jobs

More Jobs at Core Specialty

More Enterprise Technology Jobs

Find similar Site Reliability Engineer - Platform jobs: