Jobiglo

No results.

Platform Engineer - AI Infrastructure

sarvam · Bengaluru

New
Senior 🇬🇧 English
Go Python Kubernetes GPU scheduling MIG RDMA Terraform Crossplane Kueue Volcano Slurm Run:ai

Job description

About the role

Sarvam operates a large, multi‑vendor GPU fleet that simultaneously runs massive training jobs and latency‑critical inference services. This role builds the platform that sits on top of that fleet – the scheduling, scaling, multi‑tenancy, serving, and self‑service layers that let ML teams use thousands of GPUs without manual intervention. It focuses on building control‑plane services rather than operating the fleet.

Key responsibilities

  • Design and ship the serving platform that turns model artifacts into scalable, multi‑tenant endpoints with intelligent routing, rollout, and traffic‑splitting.
  • Implement autoscaling and elasticity for both training (gang scaling, scale‑to‑fit) and inference (queue‑depth, utilization‑driven) across GPU clusters.
  • Develop scheduler integrations (Kueue, Volcano, Slurm‑on‑Kubernetes or custom controllers) with gang scheduling, priority, quota enforcement, and topology‑aware placement.
  • Build multi‑tenant isolation mechanisms, RBAC, quota systems, and audit logging for secure self‑service GPU sharing.
  • Create observability, logging, tracing, and cost‑analysis tooling that operates at fleet scale.

Required profile

  • 5+ years of experience building infrastructure or platform software, delivering services that others rely on.
  • Strong software engineering skills in Go or Python with a focus on maintainable, production‑ready systems.
  • Deep knowledge of Kubernetes internals, having written operators or controllers and understood the scheduler API.
  • Working literacy of GPU‑specific constraints such as MIG, GPU sharing, gang scheduling, topology‑aware placement, and RDMA.
  • Product mindset toward internal users, designing APIs and abstractions that drive adoption and self‑service.

Required skills

  • Go
  • Python
  • Kubernetes (controllers, operators, scheduler internals)
  • GPU scheduling concepts (MIG, gang scheduling, topology‑aware placement, RDMA)
  • Infrastructure‑as‑code tools (Terraform, Crossplane)
  • Experience with scheduling systems (Kueue, Volcano, Slurm, Run:ai)

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec sarvam.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.
Source : ats:ashby

Why are you reporting this job?

Thank you for your report. We will review this job.

Apply in 30 seconds

Enter your email to apply. An account will be created automatically.

By continuing, you accept our terms of use.

Already have an account? Login

💬 Chat with us on Telegram Chat on WhatsApp

Published 10 hours ago

Expires 1 month from now

4 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

sarvam

Bengaluru