Member of Technical Staff, Architecture & Scaling

Hark

$180K — $450K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience training large models with real cost implications on compute allocation.
  • Proven ability to design empirical experiments that test competing hypotheses.
  • Expertise in scaling laws to leverage small-scale data for large-scale applications.
  • Deep knowledge of modern training stacks and distributed training over large GPU clusters.
  • Strong engineering skills, capable of writing code and analyzing performance metrics.
  • A history of delivering contributions in real-world models, evidenced by publications or production systems.

Responsibilities

  • Conduct research on model architecture and scaling for improved efficiency.
  • Establish benchmarks and run experiments to validate scalable ideas.
  • Define training parameters at scale, including learning rates and resource allocation.
  • Investigate and implement new model architectures like mixture-of-experts and hybrid attention.
  • Identify and troubleshoot instability in large model runs, addressing loss spikes and system failures.
  • Work with multimodal models, integrating various types of data from the outset.

Benefits

  • Opportunity to work on cutting-edge research with direct production impact.
  • Access to a high-performance computing environment for large model training.
  • Collaboration with interdisciplinary teams across multimodal projects.
  • Opportunity to publish findings in influential research platforms.
  • Flexibility to define innovative training methodologies and architectures.
Full Job Description
About the Role

You'll work on the architecture and scaling of our largest models: what we train, how we train it, and how to spend the next order of magnitude of compute well. This is empirical research with a direct line to production. The recipes you set are the recipes our frontier runs use.

Responsibilities
  • Conduct research on model architecture, optimization, and scaling to improve the capability and efficiency of our largest models.
  • Establish strong baselines and run controlled experiments to determine which ideas actually scale to frontier training runs.
  • Set training recipes at scale: learning-rate schedules, context length, token budgets, and compute allocation.
  • Explore new architectures, including mixture-of-experts, hybrid attention, and long-context extension.
  • Diagnose instability in large runs: loss spikes, divergence, numerical issues, and the infrastructure failures that look like research problems.
  • Work across modalities, since our models are multimodal from pretraining forward.

Requirements
  • Hands-on experience training large models, at a scale where compute allocation and stability decisions carry real cost.
  • Strong empirical instincts: you design the experiment that distinguishes between two hypotheses instead of the one that confirms the first.
  • Fluency with scaling laws and how to use small-scale results to make a frontier-scale bet.
  • Deep familiarity with a modern training stack and distributed training across large GPU clusters.
  • Strong engineering. Research here means writing the code and reading the profiler, not handing off a spec.
  • A record of work that shipped into real models, whether that shows up as papers, systems, or production runs.

Bonus Qualifications
  • Experience with mixture-of-experts routing, sparse architectures, or long-context methods.
  • Work on data mixtures, curriculum, or tokenizer design.
  • Kernel-level optimization or mixed-precision training experience.
  • Experience with efficiency work aimed at constrained inference targets, including on-device.

Compensation

The US base salary range for this full-time position is between $180,000 - $450,000 annually.

The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience. The total compensation package may also include additional components/benefits depending on the specific role. This information will be shared if an employment offer is extended.

Similar Jobs

More Jobs at Hark

More Information Technology Jobs

Find similar Member of Technical Staff, Architecture & Scaling jobs: