Platform Engineer - AI Infrastructure
sarvam · Bengaluru
Job description
About the role
Sarvam operates a large, multi‑vendor GPU fleet that simultaneously runs massive training jobs and latency‑critical inference services. This role builds the platform that sits on top of that fleet – the scheduling, scaling, multi‑tenancy, serving, and self‑service layers that let ML teams use thousands of GPUs without manual intervention. It focuses on building control‑plane services rather than operating the fleet.
Key responsibilities
- Design and ship the serving platform that turns model artifacts into scalable, multi‑tenant endpoints with intelligent routing, rollout, and traffic‑splitting.
- Implement autoscaling and elasticity for both training (gang scaling, scale‑to‑fit) and inference (queue‑depth, utilization‑driven) across GPU clusters.
- Develop scheduler integrations (Kueue, Volcano, Slurm‑on‑Kubernetes or custom controllers) with gang scheduling, priority, quota enforcement, and topology‑aware placement.
- Build multi‑tenant isolation mechanisms, RBAC, quota systems, and audit logging for secure self‑service GPU sharing.
- Create observability, logging, tracing, and cost‑analysis tooling that operates at fleet scale.
Required profile
- 5+ years of experience building infrastructure or platform software, delivering services that others rely on.
- Strong software engineering skills in Go or Python with a focus on maintainable, production‑ready systems.
- Deep knowledge of Kubernetes internals, having written operators or controllers and understood the scheduler API.
- Working literacy of GPU‑specific constraints such as MIG, GPU sharing, gang scheduling, topology‑aware placement, and RDMA.
- Product mindset toward internal users, designing APIs and abstractions that drive adoption and self‑service.
Required skills
- Go
- Python
- Kubernetes (controllers, operators, scheduler internals)
- GPU scheduling concepts (MIG, gang scheduling, topology‑aware placement, RDMA)
- Infrastructure‑as‑code tools (Terraform, Crossplane)
- Experience with scheduling systems (Kueue, Volcano, Slurm, Run:ai)
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in India.
Salaries by job title
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
Published 10 hours ago
Expires 1 month from now
4 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
sarvam
Bengaluru
Related job offers
-
QA Tech Lead – Strategic Quality Engineering Lead
atlas Bengaluru -
SharePoint Architect – Financial Services (AI-First M365 & Power Platform)
atlas Bengaluru -
Service Desk Technical Expert L3/L4 – Unified Operations
atlas Bengaluru -
Techno Functional Consultant – SAP S/4HANA EAM & Cloud Solutions
BASF Hyderabad -
Senior Service Desk Engineer – Unified Operations
atlas Bengaluru