OverviewJob Purpose
The AI Platform Engineer is responsible for the technical implementation, maintenance, and optimization of AI/ML infrastructure. This hands-on role focuses on GPU cluster deployment, container image management, platform tooling development, and deep technical troubleshooting. In addition, the engineer deploys and maintains AI-enabled workflow automation tools across LLM, MCP, and agentic capabilities, ensuring these systems operate efficiently and securely within a containerized architecture. This includes deploying and maintaining vector store infrastructure, implementing end-to-end RAG workflows, tuning agent memory systems, and hosting and managing MCP servers. The engineer also deploys and operates Agentic AI systems, including multi-agent orchestration frameworks and tool-use pipelines. The engineer serves as a core technical contributor on the AI Platform Operations team, translating architectural decisions into working infrastructure and enabling advanced, automated workflows across the platform.
Responsibilities
- Deploy, configure, and maintain GPU clusters and associated infrastructure
- Designing, building, and maintaining the workflow automation platform that uses AI capabilities (LLM/MCP/Agentic capabilities)
- Manage NVIDIA driver versions, CUDA toolkits, and container runtimes
- Build and maintain approved container images with ML frameworks (PyTorch, TensorFlow, etc.)
- Implement monitoring, alerting, and observability for GPU infrastructure
- Deploy and maintain vector store infrastructure for RAG pipelines, agent memory, and semantic search
- Implement and maintain end-to-end RAG workflows, including document ingestion, chunking, embedding generation, and retrieval optimization
- Maintain and tune agent memory systems, including short-term context windows, long-term persistent memory stores, and episodic memory retrieval patterns
- Deploy, operate, and maintain Agentic AI systems, including multi-agent orchestration frameworks and tool-use pipelines
- Deploy, host, and maintain MCP servers within the containerized platform infrastructure
- Manage MCP server configurations, versioning, access controls, and integration with agentic workflows
- Monitor MCP server health, performance, and availability; respond to incidents and perform root cause analysis
- Develop automation and tooling to improve platform reliability and efficiency
- Provide L2/L3 technical support and vendor escalation for complex issues
- Implement security controls including network policies, RBAC, and secrets management
- Execute change requests and maintain technical documentation
- Respond to and assist in production operations in a 24/7 environment
- Provide technical analysis, resolve problems, and propose solutions
- Provide support to, and coordinate with, developers, operations staff, release engineers, and end-users
- Educate and mentor team members and operations staff
- Participate in a weekly on-call rotation for after-hours support
Knowledge and Experience
- Bachelor's degree preferred.
- 3+ years in infrastructure engineering, systems administration, or DevOps
- 3+ years in scripting and automation skills (Python, Ansible, GitOps)
- 3+ years hands-on experience with Kubernetes in production
- 2+ years experience with Linux administration
- Direct experience with GPU infrastructure (NVIDIA preferred)
- 1+ years experience using CUDA
- 1+ years experience using MCPs
- 1+ years experience with vector databases and embedding infrastructure
- 1+ years experience with RAG pipeline design and deployment
- 1+ years experience with agent memory patterns (in-context, external stores, retrieval-augmented memory)
- 1+ years experience with agentic AI systems using orchestration frameworks
- 1+ years experience with semantic search, embedding models, and ANN search techniques
- 1+ years working with workflow/orchestrion automation tools
- Experience with enterprise monitoring and observability tools
- Ability to work in a service-oriented team environment
- Project Management, organization, and time management
- Customer focused, and dedicated to the best possible user experience
- Communicate effectively with both technical and business resources
- Fluent speaking, reading, and writing in English
Desired Knowledge and Experience
- 1+ years of experience with AI developer toolkits (NVIDIA drivers, CUDA, cuDNN, and NCCL)
- 1+ years of experience with Run:AI, NVIDIA AI Enterprise, or DGX systems
- 1+ years of experience with n8n
- 1+ years of experience with GitHub Actions
#LI-SH3
#LI-ONSITE