Senior Manager, Site Reliability Engineering (SRE)
nium · Bangalore
Job description
About the role
Nium is seeking a Senior Manager of Site Reliability Engineering to lead the teams that ensure the availability, performance, and scalability of its global payments platform. This hands‑on leadership role will define the reliability roadmap, partner with product, security and infrastructure groups, and drive operational excellence across a regulated financial platform serving 100+ markets.
Key responsibilities
- Lead, mentor and grow a distributed team of SREs, setting goals, career paths and performance expectations.
- Define and drive SLIs, SLOs and error budgets for critical payment, card issuance and compliance services.
- Own end‑to‑end incident management, on‑call structures, escalation paths and blameless post‑mortems.
- Embed reliability, scalability and production readiness into the software development lifecycle.
- Build and scale observability (metrics, logging, tracing, alerting) to detect issues before impact.
- Conduct capacity planning and performance engineering for high‑volume, real‑time financial transactions.
- Champion automation, self‑healing systems and reduce toil through tooling and scripts.
- Manage disaster recovery, business continuity and chaos engineering exercises.
- Collaborate with Security and Compliance to meet PCI‑DSS, SOC 2, ISO 27001 and regional regulations.
- Represent SRE in executive reviews, translating technical risk into business‑relevant reporting.
Required profile
- 10+ years of software engineering, infrastructure or SRE experience, with at least 4 years in a people‑management or technical leadership role.
- Proven track record operating and scaling high‑availability, transaction‑heavy platforms, preferably in fintech, payments or e‑commerce.
- Strong communication skills to convey reliability concepts to engineers, product leaders and executives.
- Pragmatic, metrics‑driven approach to engineering decisions.
Required skills
- AWS cloud infrastructure.
- Kubernetes and container orchestration.
- Infrastructure‑as‑code tools such as Terraform or CloudFormation.
- Observability stacks (Prometheus, Grafana, Datadog, ELK/OpenSearch).
- Defining and operationalizing SLOs/error budgets.
- Distributed systems, databases, caching, messaging queues and microservice architectures.
- Incident response and blameless post‑mortem processes.
- Security and compliance frameworks (PCI‑DSS, SOC 2, ISO 27001).
- Scripting/automation with Python, Go or Bash.
- CI/CD pipelines (Jenkins, ArgoCD, GitLab CI).
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in India.
Salaries by job title
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
Published 9 hours ago
Expires 1 month from now
9 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
nium
Bangalore
Related job offers
-
Principal R&D Technologist – Product Development
aveva Bangalore -
Principal ERP Business Analyst (SAP)
alcon Bangalore -
Principal Software Engineer I – Full-Stack & AWS
alcon Bangalore -
Network Security Senior Manager – Global Team Lead
3m BANGALORE -
Site Reliability Engineer II (SRE II)
mastercard Pune