AI Systems Engineer

Transluce

$350K — $500K+*
Education, Government & Non-Profit
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Fluent programming skills in Python
  • Expertise in bare metal optimization and GPU acceleration
  • Experience in engineering large-scale distributed systems
  • Strong leader in maintaining code quality and complexity management
  • Bonus: Experience in setting up LLM pipelines
  • Bonus: Background in open-source community management

Responsibilities

  • Establish the overall code culture and tooling for the organization
  • Address core technical challenges across various verticals
  • Create high-concurrency container-based evaluations for rapid iteration
  • Develop deterministic sandbox execution for efficient state restoration
  • Design inference stacks that enable sophisticated model inspection
  • Implement distributed RL training for extensive concurrent rollouts
  • Build internal tools to enhance team productivity

Benefits

  • Opportunity to innovate in a mission-focused non-profit
  • Work on influential systems used by governments for AI policy
  • Collaborative culture with a focus on high-impact projects
  • Room for personal and professional growth in a fast-paced environment
  • Potential to contribute to open-source tools for community use
  • Possibility of visa sponsorship for international candidates
Full Job Description
Salary range: $350,000 - $600,000/year + benefits

About the role: We are looking for an exceptional AI systems engineer to lead the design and development of our core ML stack, building systems that can scale to thousands of GPUs and performantly query trillion-token databases.

As an early member of a highly collaborative team, you will be free to innovate and move fast, building high-impact systems from the ground up. As part of a mission-focused non-profit, your work will have high direct impact (e.g. used by governments to inform AI policy) and cross-organisational reach (open-source tools the entire community can build on).

Core responsibilities:

  • Set overall code culture and tooling for a fast-growing org
  • Help to solve our core technical challenges across verticals. Examples include:
    • Docent:
      • High-concurrency container-based evals with quick ability to iterate on interventions to agentic trajectories
      • Deterministic sandbox execution of code that can efficiently restore state from checkpoints
    • Interpretability:
      • Inference stacks that are as performant as vLLM but flexible enough to allow complex model introspection and intervention, steering, configurable sampling, etc., and that can scale to 400B+ parameter models
    • Behavior elicitation:
      • Distributed RL training and roll-outs allowing thousands of concurrent rollouts across machines
    • Build great internal tools to speed up the team
  • Help tone-set in the organization around best practices for building and path-set on what infra we should build
  • Help other team members think through infra challenges


Qualities of a strong candidate:

  • Exceptional programmer fluent in Python
  • Bare metal optimization: know GPUs, other accelerators in and out (low-level performance + optimization + parallel programming)
  • Experience engineering at scale (distributed systems, reliability, architecture design)
  • Leader on global code quality and health (designing good primitives, managing complexity and scale)
  • Bonus: can set up LLM pipelines, e.g. multiple specialized LLMs interacting with each other in a performant and reliable way
  • Bonus: experience with open-source community management


We are located in San Francisco and enthusiastic to work together in-person. We are open to sponsoring international visas.

Similar Jobs

More Jobs at Transluce

More Education, Government & Non-Profit Jobs

Find similar AI Systems Engineer jobs: