📢 New: get today's jobs on our WhatsApp Channel
Jobiglo

No results.

Senior Member of Technical Staff: ML Systems and Infrastructure

DevRev · Bangalore

Senior 🇬🇧 English
Kubernetes Helm ArgoCD Argo Workflows vLLM SGLang Triton Inference Server Ray Serve Python Go PyTorch Jax TensorFlow Prometheus Grafana OpenTelemetry CUDA

Job description

About the role

DevRev is building the future of work with Computer – an AI teammate that unifies data, tools, and workflows into an AI‑ready platform. As a Senior Member of Technical Staff you will design, build, and own the end‑to‑end platform that supports the full lifecycle of machine‑learning models, from massive‑scale distributed training to ultra‑low‑latency, highly‑available inference.

Key responsibilities

  • Architect the AI infrastructure platform, covering everything from large‑scale training pipelines to production inference services.
  • Implement and scale inference stacks for large language models using frameworks such as vLLM, TensorRT‑LLM, or SGLang, addressing throughput, latency, token streaming and automated scaling.
  • Partner with AI Research and Data Science teams to create a seamless developer experience for experiment, fine‑tuning and rapid deployment of new models.
  • Build robust CI/CD/CT pipelines with Argo Workflows, ArgoCD and GitHub Actions to automate model validation, deployment and lifecycle management.

Required profile

  • 5+ years of infrastructure or software engineering experience, with at least 2 years focused on MLOps or large‑scale ML infrastructure.
  • Bachelor’s or Master’s degree in Computer Science, Engineering or a related field.
  • Deep hands‑on expertise with Kubernetes in production and fluency with Helm, ArgoCD and Argo Workflows.
  • Strong knowledge of GPU resource management, performance optimisation and cloud‑native scaling.
  • Experience with modern LLM serving frameworks (vLLM, SGLang, Triton Inference Server, Ray Serve).
  • Proficiency in Python or Go and familiarity with PyTorch, Jax or TensorFlow.
  • Observability mindset using tools such as Prometheus, Grafana and OpenTelemetry.

Required skills

  • Kubernetes
  • Helm
  • ArgoCD
  • Argo Workflows
  • GPU programming
  • vLLM
  • SGLang
  • Triton Inference Server
  • Ray Serve
  • Python
  • Go
  • PyTorch
  • Jax
  • TensorFlow
  • Prometheus
  • Grafana
  • OpenTelemetry
  • CUDA

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec DevRev.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.
Source : ats:greenhouse

Why are you reporting this job?

Thank you for your report. We will review this job.

Apply in 30 seconds

Enter your email to apply. An account will be created automatically.

Apply now →

By continuing, you accept our terms of use.

Already have an account? Login

A question about this job?

Ask it here: you will get the full job summary by e-mail, right away.

💬 Chat with us on Telegram

Published 1 week ago

Expires 1 month from now

7 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

DevRev

Bangalore