The Mathworks

Senior Cloud Reliability Solutions Engineer

The Mathworks$120K — $145K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree and 6 years of professional experience (or equivalent qualifications)
  • Strong expertise in Azure platform services
  • Knowledge of multiple cloud environments: AWS and GCP
  • Proficiency in Infrastructure as Code and automation tools
  • Experience with reliability engineering concepts
  • Familiarity with containerization and Kubernetes
  • Understanding of enterprise infrastructure and VMware management
  • Background in security and governance best practices
  • Skilled in observability and monitoring tools
  • Excellent communication and consulting skills

Responsibilities

  • Design and deliver resilient hybrid-cloud solutions across diverse platforms
  • Partner with internal teams to evaluate requirements and recommend reliable solutions
  • Automate infrastructure processes using tools like Terraform and PowerShell
  • Implement reliability engineering practices for improved service uptime
  • Create dashboards and alerts to enhance observability
  • Ensure secure platform operations through governance and access management
  • Contribute to container and Kubernetes services
  • Support traditional and cloud-native infrastructure integration
  • Collaborate with teams to share knowledge and mentor peers
  • Engage in production support responsibilities, including on-call duties

Benefits

  • Opportunities for professional growth and development
  • Collaborative work environment
  • Access to leading-edge technology and resources
  • Flexible work arrangements
  • Health and wellness programs
Full Job Description
Job Summary

The SSG Hosting team provides reliable, secure, and scalable infrastructure services that support teams across the company. We are looking for a friendly, curious, and technically strong Senior Cloud Reliability Solutions Engineer who enjoys designing practical solutions, automating repeatable work, and partnering with application, security, finance, and infrastructure teams to make our platforms easier to use and operate. In this role, you will help shape Azure-centered solutions that support workloads across Azure, AWS, GCP, and on-premises VMware environments, while advancing reliability engineering, automation, observability, and operational excellence practices.

Responsibilities

  • Design and deliver resilient hybrid-cloud solutions across Azure, AWS, GCP, and on-premises VMware environments, with deep emphasis on Azure architecture and operations.
  • Partner with internal customers to understand requirements, evaluate tradeoffs, estimate cost and effort, and recommend solutions that are reliable, secure, supportable, and aligned with business needs.
  • Build and improve infrastructure automation using Terraform, Packer, PowerShell, Python, Git-based workflows, and CI/CD pipelines to reduce manual work and improve consistency.
  • Apply reliability engineering practices including service ownership, operational readiness, incident response, root cause analysis, capacity planning, disaster recovery, and continuous improvement.
  • Create and maintain observability through useful dashboards, actionable alerts, telemetry standards, health checks, and reliability reporting across cloud and data center platforms.
  • Support secure and governed platform operations through RBAC, least privilege access, secrets management, policy enforcement, patching practices, vulnerability remediation, tagging, and cost visibility.
  • Contribute to Kubernetes and platform services including AKS, EKS, GKE, container networking, ingress, persistent storage, Helm, and GitOps-oriented operating models.
  • Bridge traditional and cloud-native infrastructure by supporting VMware, Windows, Linux, storage, networking, DNS, backup, and hybrid connectivity patterns.
  • Collaborate openly and constructively with peers and partner teams, sharing knowledge, documenting designs, mentoring others, and learning from the team.
  • Participate in production support including a rotating on-call schedule for infrastructure services and critical operational issues.


Minimum Qualifications

  • A bachelor's degree and 6 years of professional work experience (or a master's degree and 3 years of professional work experience, or a PhD degree, or equivalent experience) is required.


Additional Qualifications

A strong candidate will bring a combination of the following skills and experiences:
  • Azure platform expertise: Azure networking, compute, storage, identity, Azure Arc, Azure Monitor, Log Analytics, Azure Policy, RBAC, Azure Update Manager, Key Vault, Backup, and Site Recovery.
  • Multi-cloud literacy: working knowledge of AWS services such as EC2, VPC, IAM, S3, EKS, CloudWatch, Systems Manager, Route 53, and Organizations; familiarity with GCP Compute Engine, VPC, IAM, Cloud Storage, GKE, and Cloud Logging/Monitoring.
  • Infrastructure as Code and automation: Terraform, Packer, Ansible or similar configuration tools, PowerShell, Python, YAML, GitHub or GitLab, and CI/CD pipeline practices.
  • Reliability engineering: SLA/SLO, error budgets, incident management, RCA, disaster recovery, high availability, capacity planning, and operational readiness reviews.
  • Containers and Kubernetes: Docker, Kubernetes, AKS, EKS, GKE, Helm, ingress controllers, container networking, persistent storage, secrets management, and cluster lifecycle practices.
  • VMware and enterprise infrastructure: vSphere, ESXi, vCenter, virtual networking, storage, backup, migration planning, and hybrid integration patterns.
  • Security and governance: identity and access management, least privilege, network segmentation, certificate management, vulnerability remediation, cloud policy, compliance reporting, and cost/tagging governance.
  • Observability tools: Azure Monitor, Log Analytics, Application Insights, Prometheus, Grafana, Splunk, Datadog, or similar monitoring and logging platforms.
  • Systems administration: Windows Server, Linux distributions such as Ubuntu, RHEL, or Rocky Linux, DNS, networking, patching, performance troubleshooting, and service management.
  • Communication and consulting skills: clear written documentation, design reviews, stakeholder engagement, influence without authority, and the ability to explain technical options to both engineering and business audiences.

About The Mathworks

The MathWorks, Inc. is an American software company that specializes in mathematical computing software. The company was founded in 1984 and is headquartered in Natick, Massachusetts. The MathWorks offers a range of products, including MATLAB, Simulink, and Stateflow, which are used in engineering, science, and mathematics. The company serves customers in over 100 countries and has partnerships with major technology companies such as Microsoft and Intel. In 2019, The MathWorks was named one of the best places to work by Glassdoor.
Learn more about The Mathworks
Size
5,000 employees
Industry
Founded
1984

Similar Jobs

More Jobs at The Mathworks

More Information Technology Jobs

Find similar Senior Cloud Reliability Solutions Engineer jobs: