📢 New: get today's jobs on our WhatsApp Channel
Jobiglo

No results.

This job is no longer available

This job expired on 16/09/2026. It no longer accepts applications.

Staff Site Reliability Engineer

levistraussandco · Bengaluru

🇬🇧 English
Google Cloud Platform GKE Cloud Run BigQuery Pub/Sub GCS Composer Dataflow Vertex AI Terraform Helm GitOps Observability Automated alerting Incident response Data security Encryption at rest Secrets management Audit logging Capacity planning

Job description

About the role

Levi Strauss & Co is looking for a Staff Site Reliability Engineer to join its Data & AI Platform Engineering team in Bengaluru. The role will own the reliability, scalability and operability of enterprise data and AI platforms that support product design, retail and supply‑chain operations. You will work as a hands‑on technical leader applying Google‑style SRE practices across Google Cloud Platform and Azure.

Key responsibilities

  • Define, instrument and enforce SLOs, SLIs and error budgets for all platform services.
  • Drive reduction of mean‑time‑to‑detect (MTTD) and mean‑time‑to‑recover (MTTR) through observability, automated alerting and run‑book driven incident response.
  • Lead blameless post‑mortems and turn findings into durable reliability improvements.
  • Identify, measure and eliminate operational toil, keeping it below 50 % of engineering capacity.
  • Build self‑serve infrastructure so product and data teams can safely provision, scale and operate resources.
  • Automate deployment pipelines, configuration management and operational workflows using Infrastructure‑as‑Code (Terraform, Helm, GitOps).
  • Serve as the primary GCP subject‑matter expert and guide multi‑cloud architecture decisions across GCP and Azure.
  • Design self‑healing, auto‑scaling and capacity‑planning patterns for high‑availability data and AI services.
  • Champion data security and governance, including encryption, least‑privilege IAM, secrets management and audit logging.

Required profile

  • Hands‑on technical leader with deep experience in Site Reliability Engineering.
  • Proven track record applying Google SRE principles such as toil reduction and reliability engineering.
  • Strong expertise in Google Cloud Platform services and multi‑cloud environments.
  • Ability to work collaboratively with development and operations teams to foster shared ownership.

Required skills

  • Google Cloud Platform (GKE, Cloud Run, BigQuery, Pub/Sub, GCS, Composer, Dataflow, Vertex AI)
  • Microsoft Azure
  • Infrastructure‑as‑Code: Terraform, Helm, GitOps
  • Observability, automated alerting and incident response tooling
  • Self‑service platform engineering and automation
  • Data security and governance (encryption at rest/in‑transit, IAM, secrets management, audit logging)
  • Design of self‑healing, auto‑scaling and capacity‑planning architectures

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec levistraussandco.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.
Source : ats:workday

Why are you reporting this job?

Thank you for your report. We will review this job.

A question about this job?

Ask it here: you will get the full job summary by e-mail, right away.

💬 Chat with us on Telegram

Published 2 months ago

63 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

levistraussandco

Bengaluru