AI Infrastructure Operations Engineer

Accenture

Confirmed live yesterday High trust
Remote

Quick summary

Work type
Remote
Location
Boston, MAMilwaukee, WIDallas, TXColumbus, OHKirkland, WA
Salary
$94,400–$266,300 / yr
Posted
10 days ago
Freshness
Confirmed live yesterday

Market check

Salary context

Competitive pay

How this pay compares to similar roles

Similar $174k
This role $180k
$74k most similar roles pay here $287k

This role pays more than 64% of similar roles. Most pay $144,350–$203,141 — the shaded band above. At the midpoint, this role pays about $180k versus about $174k for comparable roles.

Based on 240 similar postings.

Employer

About Accenture

Accenture is a leading global professional services company specializing in IT, strategy, consulting, and operations, with a strong focus on digital transformation, cloud computing, and artificial intelligence.

Accenture currently has 192 open roles on FindRole.

Listed pay typically runs $94,400–$266,300 across 165 roles with salary data.

Most-posted roles

View all roles at Accenture

At a glance

TL;DR · AI Infrastructure Operations Engineer

As an AI Infrastructure Operations Engineer on the Global AI Infrastructure team, you will design, build, and operate large-scale GPU and accelerated-computing infrastructure to support demanding AI training, inference, and high-performance compute workloads across cloud, on-premises, and hybrid environments. You will develop reusable operational tools, automation workflows, and platform capabilities to streamline tasks like provisioning, configuration management, and capacity planning. Your daily work involves managing GPU-based clusters in bare-metal and containerized environments using Kubernetes and Slurm. To succeed, you must possess expertise in high-bandwidth network fabrics, NVMe-oF storage architectures, and NVIDIA tools like NCCL and CUDA-X. You will utilize Python, Bash, Terraform, Ansible, and Claude Code to automate infrastructure operations. This role solves the challenge of providing resilient, scalable, and energy-efficient compute environments for complex multi-node simulation and distributed training workloads.

What you'll do

  • Design and implement accelerated-computing infrastructure solutions for AI training, inference, and high-performance workloads.
  • Deploy and manage GPU-based clusters across bare-metal and containerized environments using Kubernetes and other schedulers.
  • Build automated workflows and reusable tools for provisioning, configuration management, monitoring, and incident response.
  • Integrate infrastructure platforms with enterprise systems, security frameworks, and service-management processes.
  • Perform and automate benchmarking and validation for GPU, compute, storage, and network components.
  • Develop technical documentation including architecture diagrams, configuration baselines, and operational runbooks.
  • Provide expert troubleshooting and optimization for multi-node AI training and distributed compute workloads.

What we're looking for

  • Must have 5+ years of experience designing, deploying, and managing accelerated-computing infrastructure in on-premises, cloud, and hybrid environments.
  • Must have 5+ years of hands-on experience with GPUs, DPUs, CPUs, high-bandwidth network fabrics, and AI-based storage architectures.
  • Must have 5+ years of experience with cluster management, workload scheduling, orchestration, observability, and infrastructure automation.
  • Must have at least 6 months of experience with Claude Code, AI automation tools, Terraform, Ansible, Python, and Bash scripting.
  • Must possess a Bachelor's degree or equivalent (minimum 12 years of work experience).
  • Candidates with an Associate’s Degree must have at least 6 years of work experience.
  • Ability to travel between 25% and 60% as required by business needs and client requirements.

More like this

Similar roles

AI Infrastructure Engineer

Fortinet

New York, NY 13 days ago $215,000$350,000
Linux GPU Docker Kubernetes Python Bash CI/CD KVM FortiGate FortiManager FortiAnalyzer Monitoring Logging Alerting Networking Virtualization Performance Testing Benchmarking Automation

AI Infrastructure Engineer

Blackrock

New York, NY 65 days ago $162,000$215,000
AWS Azure GCP Terraform Ansible CloudFormation Bicep Python Java Golang CI/CD MLOps Microservices Infrastructure as Code
5+ yrs exp Hybrid

AI Infrastructure Engineer

Blackrock

New York, NY 58 days ago $162,000$215,000
AWS Azure GCP Terraform Ansible CloudFormation Bicep Python Java Golang CI/CD MLOps Microservices Infrastructure as Code
5+ yrs exp Hybrid

AI Infrastructure Engineer

Amd

San Jose, CA 6 days ago $204,000$306,000
Kubernetes Terraform GitOps ArgoCD Flux Helm Prometheus Grafana Loki CSI CNI Slurm PyTorch vLLM SGLang Infrastructure as Code
Hybrid

AI Infrastructure Software Engineer

Qualcomm

San Diego, CA 43 days ago $94,200$141,200
C C++ Python Java Embedded Systems RTOS Linux Android Windows DSP IPC Memory Management Concurrency Machine Learning Computer Vision System Optimization