About the RoleAs an AI Infrastructure Engineer, you will architect and build the virtual access interface for our physical Ai Lab, ensuring secure, scalable, and efficient remote processing capabilities. You will lead the design and implementation of infrastructure that allows BRG teams to leverage our Ai Lab's computational power remotely, while maintaining performance standards for large-scale document processing. Key responsibilities include developing customizable interfaces for different BRG groups, implementing secure access controls, and ensuring optimal resource allocation for concurrent users processing massive datasets through LLMs.
Key Responsibilities- Design and implement a virtual access layer for the physical Ai Lab infrastructure
- Build scalable remote processing capabilities supporting 100,000+ documents per day
- Create customizable, expandable interfaces for different BRG business units
- Optimize infrastructure for maximum LLM token throughput (OpenAI/Anthropic)
- Implement secure authentication and access management systems
- Ensure high availability and fault tolerance for mission-critical AI workloads
- Lead infrastructure projects from conception to production deployment
Required:- Bachelor's degree in Computer Science, Information Technology, or a related field
- Minimum six to eight (6-8) years of hands-on experience designing, deploying, and managing scalable cloud infrastructure
- Strong experience with Infrastructure as Code (IaC) tools and methodologies
- Experience designing, implementing, and maintaining scalable, secure, and cost-efficient cloud/on-prem solutions
- Proven ability to manage and lead projects to deliver high-quality, replicable solutions
- Proficiency in VCS (Git/GitHub), modern coding languages (Python, .NET, Java, etc.), Software Development Life Cycle, and CI/CD practices
- Experience with API design and implementation for distributed systems
- Knowledge of GPU infrastructure and optimization for AI workloads
- Hands-on experience with AWS Services including:
- EC2/Lambda (apps/functions)
- SageMaker (ML)
- S3 (file management)
- Fargate/ECS/EKS (containerization)
- CDK/Terraform (IaC)
- Cost Explorer/Budgets
Preferred:- Experience with LLM deployment and optimization (OpenAI, Anthropic, etc.)
- Background in building AI/ML infrastructure and platforms
- Experience with virtual desktop infrastructure (VDI) or remote access solutions
- Knowledge of distributed computing and job scheduling systems
- AWS certifications (Solutions Architect, Machine Learning, or similar)
- Experience with cost management and optimization strategies in the cloud
- Familiarity with security best practices for AI systems and data handling