Inference Engineer

Hyperbolic Labs

$150K — $180K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Strong general inference background with high-level stack understanding
  • Deep Kubernetes experience for operational production clusters
  • Understanding of key inference performance concepts
  • Familiarity with modern inference frameworks and serving engines
  • Knowledge of NVIDIA Dynamo in distributed architectures
  • Experience setting up monitoring for production inference services
  • Proven ability to build products end to end

Responsibilities

  • Build inference capabilities on top of the Forge control plane
  • Deploy and serve models across globally distributed clusters
  • Evaluate inference frameworks for production readiness
  • Setup monitoring, gateways, and endpoints for inference services
  • Optimize inference performance, including autoscaling and cache orchestration
  • Debug customer inference issues with real impact

Benefits

  • Flexible work schedule
  • Collaborative team environment
  • Opportunities for professional development
  • Access to cutting-edge technology and frameworks
  • Supportive of work-life balance
Full Job Description
About the Role

We're looking for an Inference Engineer to build inference capabilities on top of Forge, our unified control plane, so customers can consume model tokens without managing GPUs and our NeoCloud partners get a full-stack path to their own token-factory offering. You'll own how models get deployed and served across clusters distributed around the world, on heterogeneous hardware.

Deployment comes first: serving models on Forge and our Kubernetes offering, evaluating inference frameworks, and standing up the monitoring, gateways, and endpoints that make a deployment production-ready. From there the work expands into optimization, autoscaling, KV-cache orchestration, and customer inference debugging. This is the primary seat for inference at Hyperbolic - you'll build it end to end, with real influence over where the scope lands.

Who You Are
  • Strong general inference background with a broad, high-level command of the stack rather than a narrow specialty - you can reason about the whole path from request to token
  • Deep Kubernetes experience, including hands-on ability to operate clusters in production, not just deploy to them
  • Solid grasp of the concepts that govern inference performance: TTFT, disaggregated inference, speculative decoding, and KV cache and its inner workings
  • Familiarity with modern inference frameworks and serving engines, and the judgment to evaluate and select among them for a given workload
  • Working knowledge of NVIDIA Dynamo and how it fits into a distributed serving architecture
  • Experience setting up monitoring, gateways, and endpoints for production inference services
  • Proven ability to build a product end to end - you've taken something from nothing to serving real traffic
  • Strong self-initiative and comfort operating as the primary owner of an area with minimal direction
  • Generalist instincts: you're willing to pick up adjacent work when it's what the product needs


Preferred Qualifications
  • Experience spanning both inference deployment and inference optimization
  • Hands-on model optimization work - quantization, batching strategies, kernel-level tuning, or similar
  • Understanding of RDMA and high-performance networking as they apply to distributed serving
  • Experience deploying inference across heterogeneous accelerators
  • Background supporting customers directly on inference debugging and performance issues
  • Experience at a GPU cloud, inference provider, or AI infrastructure company

Similar Jobs

More Jobs at Hyperbolic Labs

  • Head of Marketing
    $160K — $200K *
    San Francisco, CA 94112 (San Francisco County)
    Enterprise Technology
    In-Person
  • Quantitative Researcher
    $150K — $180K *
    San Francisco, CA 94112 (San Francisco County)
    Finance & Insurance
    In-Person
  • Inference Engineer
    $150K — $180K *
    San Francisco, CA 94112 (San Francisco County)
    Information Technology
    In-Person
  • Capital Markets Lead
    $150K — $180K *
    San Francisco, CA 94112 (San Francisco County)
    Finance & Insurance
    In-Person
  • Supply Operations Manager
    $110K — $130K *
    San Francisco, CA 94112 (San Francisco County)
    Technical Services
    In-Person

More Information Technology Jobs

Find similar Inference Engineer jobs: