Storage Architect

Jefferson Lab

$118K — $186K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 10+ years in HPC, scientific computing, or large-scale distributed systems
  • 4+ years designing high-performance parallel or distributed file systems at petabyte scale
  • Hands-on experience in multi-tiered, resilient storage architectures across distributed locations
  • Experience developing storage performance models and capacity planning frameworks
  • Bachelor’s degree in Computer Science or relevant IT field, Master’s preferred

Responsibilities

  • Lead the design of a multi-tiered distributed storage system prioritizing reliability and performance
  • Evaluate and select storage technologies including HPC platforms and open-source file systems
  • Develop performance and cost models to predict storage capacity and ownership costs
  • Define service level objectives and metrics for storage performance
  • Establish operational framework and staffing plans for the future storage team

Benefits

  • Opportunity to lead a groundbreaking data facility project
  • Collaborative work environment with top-tier research institutions
  • Work with cutting-edge technologies tailored for modern scientific data
  • Impactful role in enhancing scientific data infrastructure
  • Potential for career growth in a leading scientific computing environment
Full Job Description
The good-faith pay range for this role is $118,400 - $186,500 per year. Actual compensation may vary and may be above the posted range based on factors such as a candidate's skills, experience, education, certifications, and work location.

What your job will be like:

Jefferson Lab, in collaboration with Lawrence Berkeley National Laboratory are building the High-Performance Data Facility (HPDF), a landmark DOE ASCR investment in scientific data infrastructure. Spanning from the Virginia coast to the California Bay Area, HPDF will deliver seamless, integrated data services across both institutions, purpose-built for the demands of modern data science and research. Unlike traditional facilities where data is an afterthought, HPDF is designed to be a data-first facility that will treat scientific data as a first-class asset and provides the capabilities researchers need to extract its full scientific value. This is a greenfield storage architecture role.

You are the principal authority on storage architecture for the High-Performance Data Facility (HPDF), leading the design of a large-scale, multi-tiered, distributed storage system purpose-built for DOE scientific data lifecycle. You research and evaluate the full landscape of commercial vendors, cloud providers, and open-source distributed storage solutions, develop performance and cost models, and work closely with peer architects to ensure coherent, well-grounded technology selections as the project advances toward its design milestones. When the facility is complete, you lead the HPDF storage operations team, owning storage infrastructure across its distributed locations.

In this job you will:
  • Design Leadership: Lead the technical design of a multi-tiered, widely distributed storage system incorporating reliability, immutability, data integrity, and high-throughput performance across distributed locations.
  • Evaluate storage technologies: HPC platforms (VAST, Weka, Spectrum Scale), open-source file systems (Lustre, DAOS, Ceph), and cloud. Engage vendors, assess solutions, and develop analyses to support architecture and acquisition decisions.
  • Performance and Cost Modeling: Develop and validate models to predict storage performance, capacity, and total cost of ownership at scale. Define KPPs and benchmarking strategies to govern storage system selection and validation.
  • Metrics and Observability: Define initial SLOs and SLIs for storage performance and availability. Architect and implement monitoring and alerting solutions supporting design validation and future facility operations.
  • Establish Storage Team Framework: Define the operational framework and staffing plans for the future storage team, which will be responsible for the lifecycle management of the entire storage infrastructure.


Lead - Supervisory - Management
  • Lead small team


Experience
  • Required: Experience with hybrid cloud architectures for HPC workloads including cloud bursting and cross-site data movement over a WAN.
  • Preferred: Experience at a DOE national laboratory, research university HPC center, or equivalent scientific computing environment.
  • Required: 10 or more years HPC, scientific computing, large-scale distributed systems, or equivalent .
  • Required: 4 or more years designing and deploying high-performance parallel or distributed file systems at petabyte scale or beyond.
  • Required: Experience designing multi-tiered, highly available, and resilient storage architectures, including data staging, archival, and immutability strategies across distributed locations.
  • Required: Hands-on experience designing and deploying parallel or distributed file systems at petabyte scale or beyond in an HPC or scientific computing environment.
  • Required: Experience developing storage performance models, capacity planning frameworks, and total cost of ownership analyses at facility scale.


Education
  • Required: Bachelor's Degree Computer Science, Information Systems, or other relevant information technology
  • Preferred: Master's Degree Computer Science, Information Systems, or other relevant information technology


Experience and Education Exchange

Education above the minimum may be substituted for experience. Relevant experience may not be substituted for education.

Knowledge, Skills, and Abilities
  • Deep expertise in HPC storage technologies including parallel file systems, distributed file systems, object storage, cloud storage, tape libraries, SAN, and NAS.
  • Deep understanding of storage networking protocols and interconnects including Fibre Channel, iSCSI, NVMe-oF, and high-speed data transfer technologies relevant to high-throughput data workflows.
  • Evaluating and deploying commercial high-performance storage solutions (e.g. VAST, Weka, IBM Spectrum Scale) and open-source distributed file systems (e.g. Ceph, Lustre, DAOS) against performance, reliability, and cost requirements.
  • Familiarity with scientific data workflows including detector data ingest, MPI-IO access patterns, and HSM. (Preffered)
  • Familiarity with data integrity, encryption, and compliance requirements including FAIR data principles, data retention policies, and relevant federal data management standards. (Preffered)

Similar Jobs

More Jobs at Jefferson Lab

More Information Technology Jobs

Find similar Storage Architect jobs: