Software Engineer, Distributed Systems

features and labels

$180K — $250K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 3+ years experience with distributed compute and orchestration platforms in Python or Rust
  • Strong grasp of distributed systems fundamentals like consensus and fault tolerance
  • Deep knowledge of computational complexity and memory allocation
  • Proven ability to design systems that scale under heavy production loads
  • Experience using observability to enhance performance and reliability
  • Excellent communication skills for cross-team technical decision-making
  • Self-starter with a focus on ownership and continuous improvement

Responsibilities

  • Build the core Python/Rust platform including request routing and AI workload orchestration
  • Design for platform evolution to accommodate a 100x traffic increase
  • Utilize AI to automate complex system-building tasks
  • Profile and optimize CPU and memory performance in low-level code

Benefits

  • Opportunity to work on interesting and challenging projects
  • Access to significant learning and growth opportunities
  • Relocation assistance to San Francisco for new hires
  • Comprehensive health, dental, and vision insurance
  • Regular team events and offsite activities
Full Job Description
About this role:

You are an experienced software engineer who thrives on building large-scale computing platforms. You have deep expertise in large scale distributed systems that deal with high complexity, a lot of traffic and data. You know how to achieve reliability and scale with minimum operational load.
Key responsibilities
  • Build our core Python/Rust platform: request routing, AI workload orchestration, scheduling, GPU autoscaling, large scale file storage, queueing, etc
  • Produce forward designs for platform evolution as we scale to 100x current traffic and need to provide low latency across the world
  • Leverage AI to an extreme level to automate the mundane parts of building complex but reliable systems
  • Profile and tune low level CPU and memory performance
Requirements
  • 3+ years experience building distributed compute and orchestration platforms in Python or Rust
  • Strong understanding of distributed systems fundamentals: consensus, scheduling, fault tolerance, capacity planning
  • Deep understanding of computational complexity and memory allocation
  • Track record of designing systems that scale under real production load
  • Experience building and using observability to drive performance and reliability decisions
  • Excellent communication and ability to drive technical decisions across teams
  • Self-starter who executes quickly, takes ownership, and constantly seeks improvement
Nice to have
  • Experience with AI/ML inference or training infrastructure
  • Experience with high-performance systems programming (async runtimes, zero-copy, memory-safe concurrency)
  • Background in building multi-tenant compute platforms
  • Understanding of networking fundamentals and performance characteristics
  • Familiarity with GPU workload characteristics and scheduling constraints
Compensation
  • $180,000-250,000 plus equity + benefits (This range is across all 3 levels Mid, Senior and Staff)
Location
  • San Francisco, CA (willing to consider remote for Senior and Staff levels)
What we offer at fal
  • Interesting and challenging work
  • A lot of learning and growth opportunities
  • We are currently hiring in downtown San Francisco.
  • We offer relocation assistance to San Francisco.
  • Health, dental, and vision insurance (US)
  • Regular team events and offsites

Similar Jobs

More Jobs at features and labels

More Information Technology Jobs

Find similar Software Engineer, Distributed Systems jobs: