VAST Data

Senior Lab Reliability Engineer

VAST Data$120K — $145K *
US-AnywhereRemote in United States
Enterprise Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 4+ years of experience in systems or storage engineering roles.
  • Hands-on experience with enterprise storage systems like VAST, Pure, or NetApp.
  • Strong skills in Linux systems administration, including networking and storage.
  • Proficient in scripting or programming languages like Python or Bash.
  • Experience with Docker and Kubernetes in operational settings.
  • Familiar with infrastructure-as-code tools like Ansible.
  • Solid knowledge of networking concepts and virtualization platforms.

Responsibilities

  • Own the reliability of VAST clusters in the lab, including monitoring and issue resolution.
  • Act as the main technical escalation point for cluster-related challenges.
  • Shape the strategy for automation and internal tooling in the lab.
  • Establish standards for infrastructure-as-code across the lab environment.
  • Reproduce complex issues in lab settings, providing diagnostic data.
  • Manage broader infrastructure components including virtualization and networking.
  • Mentor team members to enhance their technical skills.
  • Collaborate with pre-sales engineers and professional services to validate solutions.

Benefits

  • Flexible work arrangements including remote opportunities.
  • Collaborative environment with a focus on technology innovation.
  • Opportunities for professional development and mentorship.
  • Exposure to cutting-edge technologies in storage and networking.
Full Job Description
Description

This role holds the senior technical ownership on that team. You'll own the reliability of the VAST clusters in our lab environment, influence the direction for the automation, tooling, and infrastructure-as-code practices the rest of the team builds on and executes within, and serve as the escalation point for the hardest technical problems in the lab. You'll partner with a growing team of lab and platform engineers, as well as the internal teams who rely on these tools day to day, to understand their needs and raise the operational bar of everything we run.

Responsibilities

  • Own the operational reliability of VAST clusters in the lab environment, including proactive health monitoring, upgrade planning, and issue resolution
  • Serve as the primary technical escalation point for complex cluster issues, working hands-on-keyboard to resolve them and partnering with VAST engineering when deeper investigation is needed
  • Shape the automation and tooling strategy for the lab environment, including provisioning scripts, CLI utilities, monitoring dashboards, and internal tooling that the rest of the team builds on
  • Help establish standards for infrastructure-as-code and configuration management (Ansible or equivalent) across the lab environment
  • Reproduce and isolate difficult issues in controlled lab environments, producing high-quality diagnostic data and reports for engineering
  • Own broader systems infrastructure supporting the lab, including virtualization platforms (VMware vSphere, Proxmox), compute, networking, and storage
  • Mentor and support other members of the lab operations team, raising the collective technical bar
  • Partner with pre-sales SEs, professional services, and engineering to reproduce customer-relevant scenarios and validate solutions in the lab


Required Qualifications

  • 4+ years of professional experience in a systems engineering, storage engineering, customer support engineering, or related role
  • Deep hands-on experience with enterprise storage systems (VAST, Pure, NetApp, Isilon, Ceph, or similar) at an operational or reliability level
  • Strong Linux systems administration skills, including networking, storage, filesystems, systemd, and CLI tooling
  • Strong scripting/programming experience (Python, Bash, or similar), sufficient to build and maintain automation and tooling other engineers rely on
  • Experience with Docker and Kubernetes in operational environments
  • Experience with infrastructure-as-code or configuration management tools (Ansible or equivalent)
  • Solid networking foundations (VLANs, routing, subnetting, network troubleshooting)
  • Experience with virtualization platforms (VMware vSphere, ESXi, Proxmox, or equivalent)
  • Methodical approach to troubleshooting complex issues across storage, networking, and compute layers
  • Comfort working across timezones with distributed team members
  • Excellent written and verbal communication, including ability to produce clear technical documentation and reports


Preferred Qualifications

  • Existing hands-on experience with VAST Data clusters
  • Prior experience as a Customer Support Engineer at a storage or infrastructure company OR Reliability Engineer / SRE supporting a wide variety of infrastructure services
  • Familiarity with observability platforms (Grafana, Prometheus, Elasticsearch)
  • Network switch administration experience (Cumulus, Arista, Cisco)
  • Experience with high-performance computing environments or ML/AI training workloads

About VAST Data

VAST Data is a data storage company that provides an all-flash storage platform for big data and artificial intelligence workloads. The platform enables businesses to store and analyze large amounts of data in real-time, and provides high performance and scalability. VAST Data was founded in 2016 and is headquartered in New York City.
Learn more about VAST Data
Size
50 employees
Industry
Founded
2016

Similar Jobs

More Jobs at VAST Data

More Enterprise Technology Jobs

Find similar Senior Lab Reliability Engineer jobs: