Jobiglo

No results.

Performance Engineer – Inference

sarvam · Bengaluru

New
Hybrid Senior 🇬🇧 English
SGLang vLLM NVIDIA Dynamo TensorRT-LLM Speculative decoding KV cache internals NCCL C++ CUDA Nsight Systems py-spy perf

Job description

About the role

Performance Engineer, Inference is part of Sarvam's Performance Engineering team, responsible for owning the production serving path for large distributed models end‑to‑end. The role involves integrating kernel and model artifacts into a multi‑node, multi‑tenant stack and delivering latency and throughput targets.

Key responsibilities

  • Own Sarvam's production serving path for large distributed models, from integration to deployment.
  • Read and modify source code in SGLang, vLLM, NVIDIA Dynamo or TensorRT‑LLM to adapt behavior to workload needs.
  • Operate and extend a distributed‑serving stack, including disaggregated prefill‑decode, cross‑node KV/cache transfer and routing/scheduling.
  • Build and train speculative decoding models, distill draft models, and tune acceptance rates against live serving distribution.
  • Produce and defend latency and throughput metrics (TTFT, TPOT, GPU utilization, cost per million tokens) used for company planning.

Required profile

  • 5+ years of experience in ML systems with at least 2 years on inference serving at production scale.
  • Proven track record of measurable performance improvements such as throughput gains or p99 latency reductions.
  • Hands‑on experience serving 100B+ parameter models using tensor, pipeline or expert parallelism.
  • Deep expertise in distributed serving stacks, including disaggregated prefill‑decode, KV/cache transfer and cross‑node scheduling.

Required skills

  • SGLang, vLLM, NVIDIA Dynamo or TensorRT‑LLM (source‑level fluency).
  • Speculative decoding implementation and training (draft models, distillation, acceptance‑rate tuning).
  • Deep understanding of KV cache internals, block tables, copy‑on‑write and fragmentation.
  • Proficiency with TP/PP/EP parallelism, NCCL primitives and their interaction with schedulers.
  • Strong C++ and CUDA programming skills; profiling with Nsight Systems, py‑spy or perf.

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec sarvam.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.
Source : ats:ashby

Why are you reporting this job?

Thank you for your report. We will review this job.

Apply in 30 seconds

Enter your email to apply. An account will be created automatically.

By continuing, you accept our terms of use.

Already have an account? Login

💬 Chat with us on Telegram Chat on WhatsApp

Published 9 hours ago

Expires 1 month from now

5 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

sarvam

Bengaluru