Orchestration Workload Engineer

Genentech

Confirmed live today High trust

Quick summary

Work type
On-site
Location
Kaiseraugst, SwitzerlandSan Francisco, CA
Posted
1 day ago
Freshness
Confirmed live today

Market check

Salary context

How this pay compares to similar roles

Similar $183k
$124k most similar roles pay here $236k

This listing doesn't post a salary. Most similar roles pay $147,868–$217,237.

Based on 240 similar postings.

Employer

About Genentech

Genentech is a leading research-driven company dedicated to discovering and developing, manufacturing, and commercializing medicines for people with serious and life-threatening diseases.

Genentech currently has 9 open roles on FindRole.

Listed pay typically runs $149,400–$277,500 across 8 roles with salary data.

Most-posted roles

View all roles at Genentech

At a glance

TL;DR · Orchestration Workload Engineer

As a member of the Accelerated Compute Engineering team, the Orchestration Workload Engineer serves as a subject matter expert in workload orchestration, managing the scheduler tech stack across High-Performance Computing platforms. This role involves architecting and scaling SLURM environments, designing advanced configurations like topology-aware scheduling and GPU management, and integrating containerization standards using Singularity or Apptainer. The engineer will bridge traditional scientific computing with modern AI paradigms by integrating SLURM with Kubernetes and orchestration platforms like Run:ai. Key responsibilities include solving multi-tenant bottlenecks, managing GRES/TRES modeling, and implementing Infrastructure-as-Code via Ansible and Terraform. The role requires expertise in NVIDIA MIG, InfiniBand, RoCE, MPI, and NCCL to optimize multi-node CPU and GPU environments. This position addresses the technical challenge of providing reliable, high-availability compute infrastructure for research and data science workloads within a complex pharmaceutical R&D environment.

What you'll do

  • Architect, scale, and maintain SLURM Workload Manager across heterogeneous HPC and AI environments.
  • Design and tune advanced SLURM configurations including custom plugins, topology-aware scheduling, and GPU management.
  • Integrate containerization standards like Singularity and Apptainer into the SLURM environment for hybrid workloads.
  • Bridge HPC and cloud-native ecosystems by integrating SLURM with Kubernetes and other orchestration platforms.
  • Resolve complex multi-tenant bottlenecks involving GPU allocation, MPI/NCCL communication, and hardware failures.
  • Establish workload orchestration standards and architectural patterns through leadership of global cross-functional initiatives.
  • Mentor and coach junior and mid-level engineers to develop technical expertise in workload orchestration.
  • Automate scheduler deployments and telemetry pipelines using Infrastructure-as-Code tools like Ansible and Terraform.

What we're looking for

  • Bachelor’s or advanced degree in Computer Science, Applied Mathematics, Computational Engineering, or a related technical discipline.
  • Extensive systems engineering experience with specialization in workload scheduling, SLURM administration, and multi-tenant cluster optimization.
  • Subject matter expertise in architecting, scaling, and optimizing production SLURM environments including GRES/TRES modeling for GPUs.
  • Deep expertise in SlurmDBD accounting architecture, database performance management, and scheduler telemetry.
  • Hands-on experience with Kubernetes fundamentals and container runtimes like Singularity, Apptainer, Enroot, or Docker within an HPC context.
  • Familiarity with GPU scheduling (NVIDIA MIG), high-speed interconnects (InfiniBand, RoCE), and multi-node communication frameworks (MPI, NCCL).
  • Advanced proficiency with Infrastructure-as-Code tools such as Ansible and Terraform to automate deployments and telemetry pipelines.
  • Proven experience in life sciences, pharmaceutical R&D, or high-performance scientific research environments.

More like this

Similar roles

AI Infrastructure Operations Engineer

Accenture

Remote (Boston, MA) +4 20 days ago $94,400$266,300
GPU Kubernetes Slurm Run:ai Terraform Ansible Python Bash CUDA-X NCCL REST API JSON YAML NVIDIA HPC Bare-metal NVMe-oF Parallel File Systems
5+ yrs exp Remote

HPC Architect

AbbVie

Chicago, IL 12 days ago $109,500$208,500
HPC AWS Slurm CUDA Python R Apptainer Singularity Docker MPI Posit Workbench Lustre NVIDIA InfiniBand Infrastructure as Code Linux Active Directory Jupyter NFS
7+ yrs exp

AI Operations Engineer

Cisco

Raleigh, NC 2 days ago $169,800$214,800
Python LangChain LangGraph AutoGen CrewAI RAG REST GraphQL OpenAI SDK Anthropic SDK CI/CD LLMOps Vector Databases AWS GCP Azure Prompt Engineering Model Context Protocol (MCP)
7+ yrs exp Hybrid

AI Operations Engineering Technical Leader

Cisco

Remote (Milpitas, CA) 52 days ago $216,300$280,800
CI/CD Kubernetes Docker Terraform Helm GitOps Python Bash Go AWS Azure GCP Infrastructure as Code AIOps Microservices GitHub ArgoCD Jenkins
8+ yrs exp Remote