Accenture

AI Infrastructure Operations Engineer

Accenture$80K — $196K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years designing, deploying, and managing accelerated-computing infrastructure in diverse environments.
  • 5+ years of hands-on experience with GPUs, DPUs, and AI-optimized storage designs.
  • 5+ years in cluster management and workload orchestration tools like Kubernetes and Slurm.
  • At least 6 months of experience with Claude Code, Terraform, Ansible, Python, and Bash scripting.
  • Bachelor's degree or equivalent experience (12+ years), with an exception for an Associate's degree combined with 6+ years of experience.

Responsibilities

  • Design and implement GPU infrastructure solutions that meet performance and governance standards.
  • Deploy and operate GPU clusters in bare-metal and containerized setups using Kubernetes.
  • Integrate infrastructure with enterprise systems and governance controls.
  • Build and maintain operational tools and automation workflows for infrastructure management.
  • Establish repeatable processes for cluster provisioning and incident response.
  • Automate benchmarking and validate performance across multi-node workloads.
  • Develop architecture diagrams and operational documentation to support infrastructure.

Benefits

  • Comprehensive medical, dental, and vision coverage.
  • 401(k) plan with company match.
  • Bonus opportunities based on performance.
  • Generous paid holidays and time off policy.
Full Job Description
The Global AI Infrastructure team enables resilient, high-performance compute environments for strategic clients across cloud, on-premises, and hybrid deployments. We design, build, and operate large-scale GPU and accelerated-computing infrastructure that supports demanding AI training and inference, simulation, and high-performance compute workloads. Our work spans strategy, architecture, modernization, operations, governance, and continuous improvement across the infrastructure stack. We build reusable operational tools, automation workflows, and platform capabilities that make repeatable infrastructure tasks safer, faster, and more scalable. We collaborate across the technology ecosystem to harness new capabilities, drive business transformation, and deliver dependable services at scale.

Key Responsibilities:
  • Design and implement accelerated-computing infrastructure solutions aligned to system architecture, deployment roadmaps, performance, scalability, resiliency, and governance requirements.
  • Deploy, configure, and operate GPU-based clusters across bare-metal and containerized environments, using workload schedulers and Kubernetes orchestration to support AI training, inference, and high-performance compute workloads.
  • Integrate infrastructure platforms with enterprise systems, data platforms, security frameworks, service-management processes, and governance controls.
  • Design, build, and maintain reusable tools, scripts, self-service capabilities, and automation workflows for infrastructure operations, including provisioning, configuration management, validation, capacity planning, monitoring, incident management, reporting, and recurring remediation.
  • Establish repeatable operational processes for cluster provisioning, configuration management, patching, capacity planning, monitoring, incident response, and lifecycle management.
  • Perform and automate GPU, compute, storage, and network benchmarking and validation; diagnose performance issues across multi-node AI training, inference, and distributed compute workloads.
  • Develop and maintain architecture diagrams, configuration baselines, operational runbooks, and support documentation.
  • Provide technical guidance, troubleshooting, and optimization for GPU clusters supporting AI training, inference, high-performance computing, and multi-node simulation workloads, with emphasis on availability, resiliency, scalability, energy efficiency, and cost management.


Travel may be required for this role. The amount of travel will vary from 25% to 60% depending on business need and client requirements.

Required Skills and Qualifications:
  • Minimum of 5+ years of experience designing, deploying, and managing accelerated-computing infrastructure across on-premises, cloud, and hybrid environments for hyperscaler, neocloud, large enterprise, telecommunications, financial services, manufacturing, and/or retail clients.
  • Minimum of 5+ years of hands-on experience with accelerated-computing platforms, including GPUs, DPUs, and CPUs, high-bandwidth network fabrics, and AI based storage architectures such as parallel file systems, NVMe-oF, etc.
  • Minimum of 5+ years of experience with cluster management, workload scheduling, orchestration, observability, and infrastructure automation, including building operational tools and automation workflows with platforms such as Kubernetes, Slurm, Run:ai
  • Minimum 6 months hands-on experience with Claude Code, AI automation tools, Terraform, Ansible, Python, and Bash scripting.
  • Bachelor's degree or equivalent (minimum 12 years) work experience. (If Associate's Degree, must have minimum 6 years work experience)

Preferred Skills and Qualifications:
  • Experience building AI infrastructure automation and operations tools, AgenticOps practices that enable secure, automated, governed, and reproducible platform operations.
  • Experience developing reusable infrastructure code leveraging Python, platform services, and automation workflows using REST APIs, OpenAPI, JSON/YAML schemas, webhooks, and event-driven integrations.
  • Experience operating large-scale GPU clusters, including capacity management, reliability engineering, change management, and performance validation for AI training, inference, HPC, and enterprise compute workloads.
  • Experience using NVIDIA platform tools and libraries including Base Command Manager (BCM), NGC, NCCL, CUDA-X, NVAIE, Dynamo, benchmarking tools to deploy, tune, profile, and validate cluster performance for training and inference workloads.
  • Experience managing deployments of 1,000+ GPU clusters with infrastructure services enabled for AI training, inference, high-performance, and enterprise compute environments.
  • Design and build experience in running LLMs across AI Cloud platforms from CoreWeave, Nebius, and other specialty providers
  • Knowledge of model deployments, tuning, and troubleshooting performance on AI Infrastructure
  • Industry certifications in accelerated-computing infrastructure, public cloud providers, infrastructure automation, networking, or security are a plus.


Compensation at Accenture varies depending on a wide array of factors, which may include but are not limited to the specific office location, role, skill set, and level of experience. As required by local law, Accenture provides a reasonable range of compensation for roles that may be hired as set forth below.
We anticipate this job posting will be posted until 10/13/2026.

Accenture offers a market competitive suite of benefits including medical, dental, vision, life, and long-term disability coverage, a 401(k) plan, bonus opportunities, paid holidays, and paid time off. See more information on our benefits here:

U.S. Employee Benefits | Accenture

Role Location Annual Salary Range
California $94,400 to $266,300
Cleveland $87,400 to $213,000
Colorado $94,400 to $230,000
District of Columbia $100,500 to $245,000
Illinois $87,400 to $230,000
Maine $80,400 to $196,000
Maryland $94,400 to $230,000
Massachusetts $94,400 to $245,000
Minnesota $94,400 to $230,000
New York $87,400 to $266,300
New Jersey $100,500 to $266,300
Virginia $87,400 to $245,000
Washington $100,500 to $245,000

About Accenture

Accenture plc is a multinational professional services company that provides services in strategy, consulting, digital, technology, and operations. The company has more than 537,000 employees serving clients in more than 120 countries. Accenture operates across five business segments: Communications, Media & Technology; Financial Services; Health & Public Service; Products; and Resources. The company is headquartered in Dublin, Ireland, and has offices worldwide.
Learn more about Accenture
Size
624,000 employees
Market Cap
$173.8 billion
Industry
Net Income
$5.2 billion
Founded
1989
5 Year Trend
+11.2%
Revenue
$44.7 billion
NASDAQ

Similar Jobs

More Jobs at Accenture

More Information Technology Jobs

Find similar AI Infrastructure Operations Engineer jobs: