Berkeley Research Group

AI Lab Infrastructure Engineer

Berkeley Research Group$120K — $145K *
US-AnywhereRemote in United States
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science, Information Technology, or related field
  • 4-8 years of hands-on experience with scalable cloud infrastructure
  • Strong experience with Infrastructure as Code (IaC) tools
  • Proven project management skills for high-quality solutions
  • Proficiency in VCS (Git/GitHub) and modern coding languages (Python, .NET, Java)
  • Experience with API design for distributed systems
  • Knowledge of GPU infrastructure for AI workloads
  • Hands-on experience with AWS services like EC2, S3, and SageMaker

Responsibilities

  • Design and implement a virtual access layer for the physical Ai Lab
  • Build scalable remote processing capabilities for 100,000+ documents daily
  • Create customizable interfaces for different business units
  • Optimize infrastructure for maximum LLM token throughput
  • Implement secure authentication and access management systems
  • Ensure high availability for mission-critical AI workloads
  • Lead infrastructure projects from conception to deployment

Benefits

  • Work with cutting-edge AI and cloud technologies
  • Opportunity to lead high-impact infrastructure projects
  • Collaborative environment with top-tier professionals
  • Focus on innovative solutions for complex challenges
  • Flexible work environment with remote access capabilities
Full Job Description
About the Role

As an AI Infrastructure Engineer, you will architect and build the virtual access interface for our physical Ai Lab, ensuring secure, scalable, and efficient remote processing capabilities. You will lead the design and implementation of infrastructure that allows BRG teams to leverage our Ai Lab's computational power remotely, while maintaining performance standards for large-scale document processing. Key responsibilities include developing customizable interfaces for different BRG groups, implementing secure access controls, and ensuring optimal resource allocation for concurrent users processing massive datasets through LLMs.

Key Responsibilities

  • Design and implement a virtual access layer for the physical Ai Lab infrastructure
  • Build scalable remote processing capabilities supporting 100,000+ documents per day
  • Create customizable, expandable interfaces for different BRG business units
  • Optimize infrastructure for maximum LLM token throughput (OpenAI/Anthropic)
  • Implement secure authentication and access management systems
  • Ensure high availability and fault tolerance for mission-critical AI workloads
  • Lead infrastructure projects from conception to production deployment


Required:
  • Bachelor's degree in Computer Science, Information Technology, or a related field
  • Minimum six to eight (6-8) years of hands-on experience designing, deploying, and managing scalable cloud infrastructure
  • Strong experience with Infrastructure as Code (IaC) tools and methodologies
  • Experience designing, implementing, and maintaining scalable, secure, and cost-efficient cloud/on-prem solutions
  • Proven ability to manage and lead projects to deliver high-quality, replicable solutions
  • Proficiency in VCS (Git/GitHub), modern coding languages (Python, .NET, Java, etc.), Software Development Life Cycle, and CI/CD practices
  • Experience with API design and implementation for distributed systems
  • Knowledge of GPU infrastructure and optimization for AI workloads
  • Hands-on experience with AWS Services including:
    • EC2/Lambda (apps/functions)
    • SageMaker (ML)
    • S3 (file management)
    • Fargate/ECS/EKS (containerization)
    • CDK/Terraform (IaC)
    • Cost Explorer/Budgets

Preferred:
  • Experience with LLM deployment and optimization (OpenAI, Anthropic, etc.)
  • Background in building AI/ML infrastructure and platforms
  • Experience with virtual desktop infrastructure (VDI) or remote access solutions
  • Knowledge of distributed computing and job scheduling systems
  • AWS certifications (Solutions Architect, Machine Learning, or similar)
  • Experience with cost management and optimization strategies in the cloud
  • Familiarity with security best practices for AI systems and data handling

Similar Jobs

More Jobs at Berkeley Research Group

More Information Technology Jobs

Find similar AI Lab Infrastructure Engineer jobs: