Senior Forward Deployed Engineer – AI Infrastructure
DigitalOcean · Bengaluru
Job description
About the role
DigitalOcean is seeking a Senior Forward Deployed Engineer to lead AI‑native infrastructure initiatives. You will embed with high‑value customers, design and optimise heterogeneous GPU clusters, and help shape the next generation of the company’s AI‑Native Cloud.
Key responsibilities
- Act as the primary technical authority on multi‑vendor AI infrastructure, co‑engineering custom GPU solutions for production workloads.
- Architect low‑latency, high‑throughput LLM serving platforms across NVIDIA CUDA and AMD ROCm using frameworks such as vLLM, TensorRT‑LLM, SGLang and TGI.
- Deploy, scale and manage resilient Kubernetes clusters (DOKS or bare‑metal) for compute‑heavy AI workloads, leveraging Ray, Slurm and KubeFlow.
- Automate multi‑vendor GPU provisioning, networking and storage with Terraform, Ansible and Helm.
- Debug complex stack issues spanning drivers, container runtimes, inter‑GPU communication (NCCL/RCCL) and high‑speed interconnects (InfiniBand, RoCE, Infinity Fabric).
- Translate common customer challenges into core platform features in partnership with product and infrastructure teams.
Required profile
- 6+ years in forward‑deployed engineering, AI infrastructure or technical consulting supporting production AI systems.
- Strong customer empathy and ability to translate complex infrastructure concepts into actionable solutions.
- Builder mentality focused on delivering production‑ready code, container images and deployment blueprints.
- Experience collaborating with GPU vendors, model providers or ecosystem partners on benchmarking and launch readiness.
Required skills
- Linux systems engineering
- Kubernetes and Infrastructure as Code (Terraform, Helm, Ansible)
- NVIDIA stack: CUDA, NCCL, NVLink, Triton Inference Server
- AMD stack: ROCm, HIP, RCCL, CDNA
- LLM serving frameworks: vLLM, TensorRT‑LLM, SGLang, TGI, Ray Serve
- High‑performance networking: InfiniBand, RoCE, AMD Infinity Fabric
- Storage systems: Ceph, Lustre, NVMe‑oF
- Programming: Python, Go (C++, CUDA C/C++, AMD HIP a plus)
What we offer
- Career development resources, conference reimbursement and LinkedIn Learning access.
- Competitive benefits, flexible time‑off policy and employee assistance programs.
- Performance‑based bonus, equity grants and participation in an employee stock purchase program.
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in India.
Salaries by job title
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
Published 16 hours ago
Expires 1 month from now
5 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
DigitalOcean
Bengaluru
Related job offers
-
Senior Software Engineer II – Search & AI Platform
relx Bengaluru -
Senior Software Engineer II – Search and AI Platform
relx Bengaluru -
Senior React/Next.js Developer – Contentful CMS
NexionPro Services Bengaluru -
Consulting Principal Software Engineer
relx Chennai -
Systems Engineer III – Cloud Infrastructure
relx Chennai