AI Infrastructure Engineer

Propio Language Services

• $110K — $130K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 3+ years of experience in AI/ML infrastructure, inference platforms, or distributed systems in production.
  • Hands-on production experience with AWS services like EKS, EC2 GPU workloads, and IAM/KMS.
  • Familiarity with inference stacks such as vLLM, Triton, or TensorRT-LLM.
  • Experience in operating ML/LLM systems, covering model serving and performance benchmarking.
  • Understanding of LLM Ops practices and monitoring methods.

Responsibilities

  • Design and operate low-latency streaming inference systems on AWS.
  • Support the research team's AWS-based training environment.
  • Build model registries and automated deployment pipelines for AI models.
  • Partner with teams to create workflows for edge AI model deployment.
  • Manage production readiness, reliability, and incident response for the AI platform.

Benefits

  • Opportunities for continuous learning and skill development.
  • Support for personal growth and career progression.
  • Encouragement to bring innovative ideas and challenge norms.
  • Direct impact on the company's success and mission.
Full Job Description
Job Type

Full-time

Description

AI Infrastructure Engineer

As an AI Infrastructure Engineer, you will primarily design, build, and operate our AWS-based, GPU-accelerated inference and streaming serving platform, optimizing it for low latency, high concurrency, reliability, and cost. You will also support the research team's training environment, further the AI team's ML/LLMOps capabilities, and enable the development of edge AI deployments.

You'll be empowered to:
  • Take ownership of important initiatives and outcomes.
  • Drive meaningful business results.
  • Influence decisions and contribute new ideas.
  • Partner with talented, high-performing team members.
  • Challenge yourself through continuous learning and growth.
  • Help shape the future of a rapidly growing organization


What You'll Own
  • Real-time inference and serving. Design, deploy, and operate low-latency streaming inference systems on AWS for LLM, ASR, TTS, and multimodal models. Optimize time to first token/audio, p95/p99 end-to-end latency, real-time factor, throughput, concurrency, GPU utilization, and cost per stream using technologies such as vLLM, SGLang, TensorRT-LLM, Triton, TensorRT.
  • Training environment. Support and evolve the research team's AWS-based training environment, including reproducible containers, GPU job scheduling, distributed job execution, checkpoint and resume capabilities, experiment tracking, model and data artifact access, and researcher self-service. Enable workloads using FSDP, DeepSpeed, Megatron-LM, Slurm, EKS, or SageMaker.
  • ML/LLMOps. Build model and artifact registries, lineage and versioning, automated evaluation gates, deployment pipelines, shadow and canary releases, rollback workflows, runtime and configuration management, and production observability
  • Edge AI infrastructure. Partner with researchers and device and embedded teams to build repeatable model optimization, packaging, validation, and deployment workflows for resource-constrained edge targets. Support model export and compilation, post-training quantization, runtime integration.
  • Reliability and security. Own capacity planning, production readiness, incident response, disaster recovery, and the secure operation of the AI platform. Apply least-privilege IAM, KMS encryption, private networking, secrets management.


What Makes Someone Successful Here

The most successful people at Propio aren't necessarily the ones with the longest resumes. They're the people who:
  • Take ownership instead of waiting for direction.
  • Embrace challenges as opportunities to grow.
  • Continuously seek better ways of working.
  • Turn ideas into action.
  • Hold themselves and others accountable to high standards.
  • Are driven by making a measureable impact.


Requirements

What You'll Bring

Required Qualifications:
  • 3+ years of experience building and operating AI/ML infrastructure, inference platforms, or distributed systems in production.
  • Hands-on production experience with AWS, particularly EKS, EC2 GPU workloads, ECR, S3, IAM/KMS, VPC networking, and CloudWatch and/or OpenTelemetry.
  • Hands-on experience with at least one inference stack, such as vLLM, SGLang, TensorRT-LLM, Triton, KServe, or Ray Serve.
  • Experience operating ML/LLM systems in production, including model serving, autoscaling, monitoring, incident response, and performance benchmarking.
  • Familiarity with GPU infrastructure and at least one serving stack, such as vLLM, SGLang, TensorRT-LLM, or Triton.
  • A working understanding of LLM Ops practices, including evaluation, observability and tracing, cost control, and versioning.

Preferred Qualifications:
  • Low-latency, real-time, or streaming inference experience, especially for audio/speech-directly relevant to interpretation.
  • Experience with real-time audio pipelines and streaming protocols, including WebRTC, WebSocket, and gRPC streaming.
  • Experience supporting distributed training environments using FSDP, DeepSpeed, Megatron-LM, Slurm, EKS, or SageMaker HyperPod.
  • Inference optimization experience, including quantization, KV-cache optimization, speculative decoding, and continuous batching.
  • Experience building Edge AI deployment toolchains using ONNX Runtime, TensorRT/Jetson, ExecuTorch, llama.cpp, MLC-LLM, or similar runtimes.
  • Experience with SRE practices, capacity planning, production incident response, and secure infrastructure for PHI/PII.
  • Contributions to relevant open-source infrastructure, serving, observability, or edge-runtime projects.


Even if your experience doesn't perfectly match every qualification, we encourage you to apply. We're looking for potential, drive, and a commitment to growth as much as experience.

What You'll Gain

Own Your Growth: We invest in people who invest in themselves. You'll have opportunities to learn, develop, and expand your capabilities while building a meaningful career.

Own Your Impact: You'll see the connection between your work and our success. We believe great people deserve the opportunity to make a real difference.

Own Your Innovation: The best ideas can come from anywhere. We encourage curiosity, creativity, and challenging the way things have always been done.

Own Your Success: Whether you're building expertise, pursuing leadership opportunities, or expanding your career path, we'll give you room to grow and the support to get there.

At Propio, your work doesn't just move a company forward. It helps connect people, communities, and opportunities across the world. Ready to Build Something Bigger? Apply today and discover what happens when you own your success.

#LI-JS1

Similar Jobs

More Jobs at Propio Language Services

More Information Technology Jobs

Find similar AI Infrastructure Engineer jobs: