Member of Technical Staff, AI Compute & Data Infrastructure

Vinci4D

$160K — $200K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 10+ years experience in large-scale distributed systems, particularly with GPU infrastructure for model training.
  • Hands-on experience with multi-node distributed training using modern hardware, including performance tuning and failure recovery.
  • Significant experience managing petabyte-scale training data, balancing throughput and storage cost.
  • Proficient in setting scheduling and quota policies for shared GPU resources, optimizing for utilization and queue times.
  • Proven track record of developing user-friendly infrastructure for researchers and engineers, focusing on user outcomes rather than system metrics.
  • Capability to adapt infrastructure to significant changes in scale or workload type, with clear rationale on modifications made.
  • Experience mentoring engineers and teaching cross-disciplinary knowledge to empower teams.

Responsibilities

  • Define and manage compute architecture for GPU infrastructure as the team and data grow.
  • Oversee workload scheduling to optimize capacity for training runs, validation, and data preparation while transitioning to algorithmic solutions as needed.
  • Design and implement data infrastructure to support petabyte-scale training jobs with updated data types and sources.
  • Manage the training and validation platform to ensure seamless operation for the AI team's projects.
  • Establish reliability and cost-efficiency metrics for compute clusters, focusing on utilization and job success rates.
  • Engage in hands-on engineering by producing and reviewing production code and troubleshooting complex issues.
  • Provide mentorship and establish best practices for a growing engineering team, ensuring institutional knowledge is documented.
  • Collaborate with leadership and the AI team on future infrastructure needs based on upcoming model requirements.

Benefits

  • Flexible work environment with opportunities for remote work.
  • The chance to shape foundational infrastructure and practices for a cutting-edge team.
  • Support for professional development and continuing education in emerging technologies.
  • Inclusive company culture that values collaboration.
  • Access to industry-leading tools and technologies for AI model training.
Full Job Description
The role

Training a foundation model for physics means holding petabytes of simulation data and keeping GPU clusters saturated with it. We are hiring the engineer who will own that layer: the storage where our training data lives, and the compute the AI team trains and validates models on.

This is a founding role for the discipline, with a wide surface and a small team. You will set the compute and data architecture, and you will also be the person who finds out why a job has been sitting in pending for two hours. The AI team decides what to train and how to judge the result. Your job is to make sure the compute and the data are there when they need them, that the runs finish, and that we can afford it. The measure of the role is what the AI team can do with what you build: how many experiments they can run, how quickly results come back, how reliably a long run finishes without an infrastructure failure, and how much training we get for what we spend.

Nothing in this layer stays fixed. Our training data will grow, and it will take on new sources and new types alongside what we have today. GPU hardware and the market for it move faster than most infrastructure, and our appetite for compute keeps increasing. We are hiring the person who leads us through that on the compute and data side: who identifies the next step before we are forced into it, makes the case for what it costs, and then builds it.

What you'll do
  • Compute architecture. Define how our GPU capacity is organized, provisioned, and grown, and how that architecture holds up as both the fleet and the team get larger.
  • Workload scheduling. Own how competing work claims that capacity: queueing, priority, preemption, gang scheduling, and quota between training runs, validation jobs, and data preparation. This starts as a judgment call among a handful of stakeholders and becomes an allocation problem worth solving algorithmically. Recognizing when that transition arrives, and building for it, is part of the role.
  • Data infrastructure at petabyte scale. Design the storage, ingestion, and access paths that keep training jobs fed at full throughput, and keep datasets versioned and reproducible as they change. Build for new sources and data types arriving at similar or larger scale, not only for what we hold today.
  • Training and validation platform. Own the systems the AI team uses to launch, checkpoint, resume, and monitor runs, and make sure validation workloads get the compute and data access they need without competing with training for it.
  • Reliability, utilization, and cost. Set and meet targets for cluster utilization, job success rate, and time from submission to first batch. Own the cost of a training run and be able to account for where it goes.
  • Hands-on engineering. Write and review production code, lead architecture reviews, and take the hardest debugging problems yourself.
  • Mentorship and teaching. Mentor the engineers who join this team as it grows, and help hire them. Review designs outside your own work. Make the AI team more capable with the infrastructure than they were before, and write things down so the answer to a recurring question lives somewhere other than in your head.
  • Technical direction. Work with the AI team on what the next model will require from infrastructure before it is required, and with leadership on capacity planning and compute investment. Say what our compute and data footprint needs to look like a year out and what it will cost. The leadership in this role is technical, not managerial.


What we're looking for
  • 10+ years building large-scale distributed systems, including several years running GPU infrastructure for large model training
  • Direct experience operating multi-node distributed training on modern accelerators: runs that held dozens or hundreds of GPUs for days at a time, with the scheduling, interconnect behavior, checkpointing, and failure recovery that requires. You have found out why a run was slower than the hardware allowed and fixed it.
  • Direct experience serving training data at petabyte scale, where throughput and storage cost were both constraints you had to answer for
  • Experience setting scheduling and quota policy on shared GPU capacity, with a view on the tradeoff between fleet utilization and how long people wait in the queue
  • A track record of building infrastructure whose users are researchers and engineers, and of being measured on what those users were able to do with it, not on the system itself
  • Infrastructure you built that survived a substantial change in scale or in the character of the workload, with a clear account of what held up and what you had to replace
  • A history of mentoring engineers and of teaching people outside your specialty enough to work on their own
  • The ability to set technical direction, make the case for it to people outside infrastructure, and then implement it. At this stage the role is one engineer and a small team, not one engineer directing several.
  • Advanced degree in computer science or a related field, or equivalent industry experience


Nice to have
  • Kubernetes, Ray, Slurm, Airflow, or comparable orchestration running ML workloads in production
  • Hybrid or multi-cloud GPU capacity, including capacity planning and cost negotiation with providers
  • Fluency in a modern training stack: PyTorch distributed, FSDP or DeepSpeed, mixed precision, GPU profiling
  • Applied optimization for scheduling or allocation problems, such as linear programming, graph algorithms, or heuristic solvers
  • Experiment tracking and model artifact management
  • Work with large numerical or geometric datasets, such as simulation, EDA, CAD, imaging, or autonomous systems


Is this you?

This role writes and owns the code. If your recent work has been directing an infrastructure organization, setting standards for other teams, or managing vendor relationships, this is probably not the right fit. If you have been the person who owns a cluster and everything that runs on it, it likely is.

Much of what a platform team would have handled for you at a larger company does not exist here yet. Building it is the job.

We are also hiring for the scale you have already worked at. This role assumes production training runs on large multi-GPU clusters, and petabyte-scale data serving under a budget. Single-node or lab-scale model work does not transfer to these problems.

Similar Jobs

More Jobs at Vinci4D

More Information Technology Jobs

Find similar Member of Technical Staff, AI Compute & Data Infrastructure jobs: