Engineer, Platform Engineering - AI

Intercontinental Exchange Holdings, Inc.

• $110K — $130K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 3+ years in infrastructure engineering, systems administration, or DevOps.
  • 3+ years of scripting and automation skills (Python, Ansible, GitOps).
  • 3+ years hands-on experience with Kubernetes in a production environment.
  • 2+ years' experience with Linux administration.
  • Direct experience with GPU infrastructure, preferably NVIDIA.

Responsibilities

  • Deploy, configure, and maintain GPU clusters and associated infrastructure.
  • Design, build, and maintain the AI workflow automation platform.
  • Manage NVIDIA driver versions, CUDA toolkits, and container runtimes.
  • Implement monitoring, alerting, and observability for GPU infrastructure.
  • Provide L2/L3 technical support and vendor escalation for complex issues.

Benefits

  • Flexible working environment with a start-up culture.
  • Opportunity to work with leading AI/ML technologies.
  • Engagement in problem-solving and decision-making processes.
  • Potential for rapid career advancement with a focus on growth.
  • Focus on teamwork and collaboration for enhancing user experience.
Full Job Description
Overview

Job Purpose

The AI Platform Engineer is responsible for the technical implementation, maintenance, and optimization of AI/ML infrastructure. This hands-on role focuses on GPU cluster deployment, container image management, platform tooling development, and deep technical troubleshooting. In addition, the engineer deploys and maintains AI-enabled workflow automation tools across LLM, MCP, and agentic capabilities, ensuring these systems operate efficiently and securely within a containerized architecture. This includes deploying and maintaining vector store infrastructure, implementing end-to-end RAG workflows, tuning agent memory systems, and hosting and managing MCP servers. The engineer also deploys and operates Agentic AI systems, including multi-agent orchestration frameworks and tool-use pipelines. The engineer serves as a core technical contributor on the AI Platform Operations team, translating architectural decisions into working infrastructure and enabling advanced, automated workflows across the platform.

 

Responsibilities

  • Deploy, configure, and maintain GPU clusters and associated infrastructure
  • Designing, building, and maintaining the workflow automation platform that uses AI capabilities (LLM/MCP/Agentic capabilities)
  • Manage NVIDIA driver versions, CUDA toolkits, and container runtimes
  • Build and maintain approved container images with ML frameworks (PyTorch, TensorFlow, etc.)
  • Implement monitoring, alerting, and observability for GPU infrastructure
  • Deploy and maintain vector store infrastructure for RAG pipelines, agent memory, and semantic search
  • Implement and maintain end-to-end RAG workflows, including document ingestion, chunking, embedding generation, and retrieval optimization
  • Maintain and tune agent memory systems, including short-term context windows, long-term persistent memory stores, and episodic memory retrieval patterns
  • Deploy, operate, and maintain Agentic AI systems, including multi-agent orchestration frameworks and tool-use pipelines
  • Deploy, host, and maintain MCP servers within the containerized platform infrastructure
  • Manage MCP server configurations, versioning, access controls, and integration with agentic workflows
  • Monitor MCP server health, performance, and availability; respond to incidents and perform root cause analysis
  • Develop automation and tooling to improve platform reliability and efficiency
  • Provide L2/L3 technical support and vendor escalation for complex issues
  • Implement security controls including network policies, RBAC, and secrets management
  • Execute change requests and maintain technical documentation
  • Respond to and assist in production operations in a 24/7 environment
  • Provide technical analysis, resolve problems, and propose solutions
  • Provide support to, and coordinate with, developers, operations staff, release engineers, and end-users
  • Educate and mentor team members and operations staff
  • Participate in a weekly on-call rotation for after-hours support

 

Knowledge and Experience

  • Bachelor's degree preferred.
  • 3+ years in infrastructure engineering, systems administration, or DevOps
  • 3+ years in scripting and automation skills (Python, Ansible, GitOps)
  • 3+ years hands-on experience with Kubernetes in production
  • 2+ years experience with Linux administration
  • Direct experience with GPU infrastructure (NVIDIA preferred)
  • 1+ years experience using CUDA
  • 1+ years experience using MCPs
  • 1+ years experience with vector databases and embedding infrastructure
  • 1+ years experience with RAG pipeline design and deployment
  • 1+ years experience with agent memory patterns (in-context, external stores, retrieval-augmented memory)
  • 1+ years experience with agentic AI systems using orchestration frameworks
  • 1+ years experience with semantic search, embedding models, and ANN search techniques
  • 1+ years working with workflow/orchestrion automation tools
  • Experience with enterprise monitoring and observability tools
  • Ability to work in a service-oriented team environment
  • Project Management, organization, and time management
  • Customer focused, and dedicated to the best possible user experience
  • Communicate effectively with both technical and business resources
  • Fluent speaking, reading, and writing in English

 

Desired Knowledge and Experience 

  • 1+ years of experience with AI developer toolkits (NVIDIA drivers, CUDA, cuDNN, and NCCL)
  • 1+ years of experience with Run:AI, NVIDIA AI Enterprise, or DGX systems
  • 1+ years of experience with n8n
  • 1+ years of experience with GitHub Actions

#LI-SH3

#LI-ONSITE

Similar Jobs

More Jobs at Intercontinental Exchange Holdings, Inc.

More Information Technology Jobs

Find similar Engineer, Platform Engineering - AI jobs: