Jobverse
DevOps Engineer ยท India

Senior Site Reliability Engineer

weekday-1ยทBengaluru, Karnataka, India

View & apply on company site โ†’

๐—ง๐—ต๐—ถ๐˜€ ๐—ฟ๐—ผ๐—น๐—ฒ ๐—ถ๐˜€ ๐—ณ๐—ผ๐—ฟ ๐—ผ๐—ป๐—ฒ ๐—ผ๐—ณ ๐˜๐—ต๐—ฒ ๐—ช๐—ฒ๐—ฒ๐—ธ๐—ฑ๐—ฎ๐˜†'๐˜€ ๐—ฐ๐—น๐—ถ๐—ฒ๐—ป๐˜๐˜€

๐—ฆ๐—ฎ๐—น๐—ฎ๐—ฟ๐˜† ๐—ฟ๐—ฎ๐—ป๐—ด๐—ฒ: ๐—ฅ๐˜€ ๐Ÿญ๐Ÿฏ๐Ÿฌ๐Ÿฌ๐Ÿฌ๐Ÿฌ๐Ÿฌ - ๐—ฅ๐˜€ ๐Ÿฎ๐Ÿฌ๐Ÿฌ๐Ÿฌ๐Ÿฌ๐Ÿฌ๐Ÿฌ (๐—ถ๐—ฒ ๐—œ๐—ก๐—ฅ ๐Ÿญ๐Ÿฏ-๐Ÿฎ๐Ÿฌ ๐—Ÿ๐—ฃ๐—”)

Experience: 4+ yrs

Location: Bengaluru, Karnataka, India

Job Type: Full-time

We are looking for an experienced Senior Site Reliability Engineer (SRE) to build, operate, and continuously improve highly reliable, scalable, secure, and high-performing production systems across hybrid and multi-cloud environments .

The role combines cloud infrastructure, Kubernetes, automation, observability, incident management, and reliability engineering. The ideal candidate will have strong hands-on experience with AWS, Kubernetes, Terraform, Python, Bash, and modern observability platforms , along with a strong understanding of production operations and distributed systems.

Requirements

Key Responsibilities

  • Define and manage SLIs, SLOs, SLAs, error budgets, and reliability objectives for critical production services.
  • Drive initiatives to improve system availability, scalability, performance, resilience, and operational efficiency.
  • Manage and support production Kubernetes environments , including Amazon EKS and Red Hat OpenShift.
  • Deploy and maintain containerised workloads using Docker, Kubernetes, and Helm .
  • Manage cloud infrastructure across AWS and IBM Cloud , including hybrid-cloud environments.
  • Design and maintain reliable cloud connectivity, networking, disaster-recovery, and failover solutions.
  • Develop and maintain infrastructure using Terraform and Infrastructure as Code (IaC) practices.
  • Automate operational processes, infrastructure tasks, and troubleshooting workflows using Python and Bash .
  • Build and enhance observability solutions using Prometheus, Grafana, OpenTelemetry, Thanos , and logging platforms.
  • Monitor system health, identify performance bottlenecks, and proactively address reliability risks.
  • Participate in and lead high-severity incident response and production troubleshooting.
  • Conduct root-cause analysis and lead post-incident reviews and corrective actions.
  • Develop and maintain capacity-planning and reliability-improvement strategies.
  • Implement secure, resilient, and compliant infrastructure practices across cloud environments.
  • Support disaster-recovery planning, testing, and continuous improvement.
  • Collaborate with software engineering, platform, security, and architecture teams to improve production reliability.
  • Contribute to architecture reviews, engineering standards, operational best practices, and automation initiatives.
  • Mentor engineers and promote strong SRE, DevOps, observability, and production-engineering practices.

What Makes You a Great Fit

  • 4โ€“6 years of professional experience in Site Reliability Engineering, DevOps, Cloud Infrastructure, or a closely related field.
  • Strong hands-on experience with AWS and Kubernetes in production environments.
  • Experience managing Amazon EKS, Docker, and Helm .
  • Practical experience with Red Hat OpenShift is highly desirable.
  • Strong proficiency in Terraform and Infrastructure as Code practices.
  • Hands-on scripting and automation experience using Python and Bash .
  • Strong experience with Prometheus and Grafana for monitoring and observability.
  • Experience with OpenTelemetry, Thanos, logging platforms , or similar observability technologies.
  • Strong understanding of SLIs, SLOs, error budgets, incident management, and production troubleshooting .
  • Good understanding of DNS, TCP/IP networking, TLS, VPNs, load balancing, firewalls, and cloud connectivity.
  • Experience working with hybrid or multi-cloud infrastructure, preferably including AWS and IBM Cloud .
  • Strong understanding of containers, distributed systems, scalability, availability, and fault tolerance.
  • Experience with disaster recovery, capacity planning, and production resilience.
  • Exposure to regulated or compliance-driven environments such as HIPAA, SOC 2, PCI DSS, or ISO 27001 .
  • Strong analytical, troubleshooting, and root-cause analysis skills.
  • Excellent communication and collaboration skills.
  • Ability to take ownership of critical production systems and operate effectively during high-severity incidents.
  • Experience mentoring engineers and contributing to technical architecture and reliability standards.
View & apply on company site โ†’

Sourced from a public career listing. Jobverse is an aggregator, not the employer.

โ† All DevOps Engineer jobs