Forward Deployment Engineer (Inference & RL POC)

Glint Tech Solutions LLC

$120K — $160K *
Enterprise Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years of software engineering experience, with Python as a core language
  • Hands-on experience with machine learning inference or training systems
  • Familiarity with distributed systems architecture and GPUs
  • Experience interacting with customers and managing ambiguous requirements
  • Proficient in debugging end-to-end systems including code, infrastructure, networking, and performance

Responsibilities

  • Design and implement POCs that integrate customer needs with machine learning systems
  • Deploy and optimize large language model (LLM) inference and reinforcement learning (RL) on GPU clusters
  • Collaborate directly with research teams and startups to refine ML systems in real-time
  • Troubleshoot performance and stability issues in practical deployment settings
  • Optimize inference stacks for performance metrics including latency and throughput
  • Conduct rigorous testing on multi-GPU and multi-node environments to assess system reliability
  • Provide feedback to enhance product offerings based on hands-on customer engagement

Benefits

  • Opportunity to work closely with real users and cutting-edge GPU technology
  • Engage in impactful work on inference and RL applications
  • Contribute directly to product development and improvements through customer insights
  • Experience fast-paced iteration and have a high degree of ownership over projects
Full Job Description
About the job Forward Deployment Engineer (Inference & RL POC)

Location: Bay area (frequent customer interaction)

Team: Inference & Reinforcement Learning Platform

About the Role

We're looking for a Forward Deployment Engineer (FDE) to work directly with customers and partners to design, deploy, and validate inference and reinforcement learning (RL) proof-of-concepts on GMI's GPU infrastructure.

This is a high-impact, hybrid engineering role that sits at the intersection of platform engineering, applied ML, and customer success. You'll be embedded with customers during early-stage deployments-turning research ideas, datasets, and business requirements into working, performant systems on real GPU clusters.

If you enjoy being close to users, debugging real systems, and shipping results fast (not just writing docs), this role is for you.

Description

Own customer POCs end-to-end
  • Deploy and optimize LLM inference, RL training, and post-training workflows on GMI clusters
  • Translate customer requirements into concrete system designs and experiments

Forward-deploy with customers
  • Work hands-on with research teams, startups, and enterprise customers
  • Debug performance, stability, and correctness issues in real environments

Inference deployment
  • Stand up and tune inference stacks (e.g. vLLM / SGLang / Ray Serve-style architectures)
  • Optimize latency, throughput, GPU utilization, and cost efficiency

RL & post-training POCs
  • Support RLHF / RFT / SFT workflows using customer-provided datasets
  • Integrate SDKs, training APIs, and cluster resources to shorten idea experiment cycles

Performance & reliability
  • Diagnose GPU, networking, and distributed system bottlenecks
  • Run benchmarks, profiling, and stress tests on multi-GPU / multi-node setups

Feedback loop to product
  • Feed real-world customer learnings back into GMI's platform, SDKs, and APIs
  • Help shape reference architectures, cookbooks, and best practices

Requirements

Core Requirements
  • Strong software engineering background (Python required; Go / Rust a plus)
  • Hands-on experience with ML inference or training systems
  • Familiarity with distributed systems and GPUs (multi-GPU, multi-node)
  • Comfort working directly with customers and ambiguous requirements
  • Ability to debug end-to-end systems (code, infra, networking, performance)

Nice to Have
  • Experience with:
  • LLM inference frameworks (vLLM, SGLang, Ray Serve, Triton, etc.)
  • RL or post-training workflows (RLHF, RFT, SFT)
  • PyTorch, DeepSpeed, Megatron-LM, or similar
  • Kubernetes-based ML platforms
  • GPU performance profiling and optimization
  • Prior experience as:
  • Forward Deployed Engineer
  • Solutions Engineer
  • ML Platform Engineer
  • Applied Research Engineer

What Makes This Role Special
  • You're close to real users and real GPUs-not abstract roadmaps
  • You'll work on cutting-edge inference and RL workloads, not toy demos
  • You'll influence product direction through direct customer feedback
  • Fast iteration, high ownership, and visible impact

Who Thrives Here
  • Engineers who like shipping over theorizing
  • People who enjoy being the last mile problem solver
  • Builders who want exposure to both deep systems and applied ML
  • Those excited by early-stage POCs that turn into real production systems

Similar Jobs

More Jobs at Glint Tech Solutions LLC

  • MicroStrategy/BI Platform Administrator
    $100K — $130K *
    Reston, VA 20191 (Fairfax County)
    Information Technology
    In-Person
  • DevOps Engineer
    $100K — $130K *
    Plano, TX 75025 (Collin County)
    Information Technology
    In-Person
  • IBM MQ Developer/Admin
    $90K — $120K *
    Deerfield Beach, FL 33442 (Broward County)
    Information Technology
    In-Person
  • DEVOPS ENGINEER
    $90K — $130K *
    Bentonville, AR 72712 (Benton County)
    Information Technology
    In-Person
  • Infrastructure Engineer
    $100K — $140K *
    Palo Alto, CA 94303 (Santa Clara County)
    Finance & Insurance
    In-Person

More Enterprise Technology Jobs

Find similar Forward Deployment Engineer (Inference & RL POC) jobs: