Jobiglo

No results.

Senior Forward Deployed Engineer – AI Infrastructure

DigitalOcean · Bengaluru

New
Hybrid Senior 🇬🇧 English
Linux Kubernetes Terraform Ansible Helm NVIDIA CUDA NCCL NVLink Triton Inference Server AMD ROCm HIP RCCL CDNA vLLM TensorRT-LLM SGLang TGI Ray Slurm KubeFlow Python Go C++ CUDA C/C++

Job description

About the role

DigitalOcean is seeking a Senior Forward Deployed Engineer to lead AI‑native infrastructure initiatives. You will embed with high‑value customers, design and optimise heterogeneous GPU clusters, and help shape the next generation of the company’s AI‑Native Cloud.

Key responsibilities

  • Act as the primary technical authority on multi‑vendor AI infrastructure, co‑engineering custom GPU solutions for production workloads.
  • Architect low‑latency, high‑throughput LLM serving platforms across NVIDIA CUDA and AMD ROCm using frameworks such as vLLM, TensorRT‑LLM, SGLang and TGI.
  • Deploy, scale and manage resilient Kubernetes clusters (DOKS or bare‑metal) for compute‑heavy AI workloads, leveraging Ray, Slurm and KubeFlow.
  • Automate multi‑vendor GPU provisioning, networking and storage with Terraform, Ansible and Helm.
  • Debug complex stack issues spanning drivers, container runtimes, inter‑GPU communication (NCCL/RCCL) and high‑speed interconnects (InfiniBand, RoCE, Infinity Fabric).
  • Translate common customer challenges into core platform features in partnership with product and infrastructure teams.

Required profile

  • 6+ years in forward‑deployed engineering, AI infrastructure or technical consulting supporting production AI systems.
  • Strong customer empathy and ability to translate complex infrastructure concepts into actionable solutions.
  • Builder mentality focused on delivering production‑ready code, container images and deployment blueprints.
  • Experience collaborating with GPU vendors, model providers or ecosystem partners on benchmarking and launch readiness.

Required skills

  • Linux systems engineering
  • Kubernetes and Infrastructure as Code (Terraform, Helm, Ansible)
  • NVIDIA stack: CUDA, NCCL, NVLink, Triton Inference Server
  • AMD stack: ROCm, HIP, RCCL, CDNA
  • LLM serving frameworks: vLLM, TensorRT‑LLM, SGLang, TGI, Ray Serve
  • High‑performance networking: InfiniBand, RoCE, AMD Infinity Fabric
  • Storage systems: Ceph, Lustre, NVMe‑oF
  • Programming: Python, Go (C++, CUDA C/C++, AMD HIP a plus)

What we offer

  • Career development resources, conference reimbursement and LinkedIn Learning access.
  • Competitive benefits, flexible time‑off policy and employee assistance programs.
  • Performance‑based bonus, equity grants and participation in an employee stock purchase program.

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec DigitalOcean.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.
Source : ats:greenhouse

Why are you reporting this job?

Thank you for your report. We will review this job.

Explore further

Salaries, guides and searches in India.

Apply in 30 seconds

Enter your email to apply. An account will be created automatically.

By continuing, you accept our terms of use.

Already have an account? Login

💬 Chat with us on Telegram Chat on WhatsApp

Published 15 hours ago

Expires 1 month from now

4 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

DigitalOcean

Bengaluru