Sr. Staff Site Reliability Engineer (Linux/Network troubleshooting/Scripting)
Zscaler · Hyderabad · 7+ yrs experience · Posted 2026-07-18
Tech stack: AWS, Ansible, C, Golang, Java, Kubernetes, Linux, Python, SQL, Terraform, Go
Apply on the company site · Get a referral for this role
Zscaler salary & ratings · Zscaler interview process · More live openings
About the role
Responsibilities:
- We are seeking an experienced Senior Staff, Site Reliability Engineer to join our dynamic SRE Cloud Infrastructure & Operations team.
- In this high-impact role, you will report directly to the Director of Site Reliability Engineering and play a pivotal part in architecting, scaling, and maintaining our next-generation cloud infrastructure.
- You will bridge the gap between development and operations, ensuring our large-scale distributed systems are highly available, secure, and incredibly resilient.
- Architecting & Automating: Design, implement, and manage various advanced cloud management automations to eliminate toil and accelerate delivery
- Container Orchestration: Oversee and optimize containerized architectures using EKS and GKE to ensure robust production performance
- Observability Systems: Lead the creation, deployment, and optimization of highly scalable monitoring and alerting systems
- Cloud Operations & Incident Management: Own cloud operations, deployments, on-call support, and incident management while continuously designing and tuning Linux and BSD-based systems
- Cross-Functional Collaboration: Serve as a core member of cross-functional project teams, contributing to technology-based solutions and consulting on concept feasibility for new initiative
Qualifications:
- AI & Automation Curiosity: Demonstrated curiosity and active exploration of AI tools, with a proven history of integrating new technologies to enhance daily workflows and augment problem-solving
- Distributed Systems Experience: Minimum of 7 years of relevant experience in designing, analyzing, and troubleshooting large-scale distributed systems
- Technical Ecosystem Mastery: Deep hands-on experience with C, Java, GoLang, Python, Terraform, Ansible, Python automation, networking, Kubernetes, and AWS cloud
- Web Protocols & Security: Comprehensive understanding of web security and core protocols including HTTP, SSL/TLS, DNS, SQL, and networking fundamentals
- Observability Architecture: Proven experience in observability, building complex dashboards, managing Grafana, and maintaining a sharp understanding of SLIs, SLOs, and error budgets
- Modern DevOps Expertise: Strong DevOps skills across CI/CD pipelines, Source Control Management (SCM), builds/releases, and Continuous Integration tools/frameworks
- Advanced knowledge of Virtualization, Cloud Architecture, modern Cloud Services, and automated deployment methodologies
- A proven track record of resolving critical escalations and proactively preventing the reoccurrence of incidents through targeted process, monitoring, and reliability improvements
- Active contribution to OS/software packaging and distribution, alongside a passion for mentoring others on SRE best practices within the team
Qualifications
- AI & Automation Curiosity:
- Demonstrated curiosity and active exploration of AI tools, with a proven history of integrating new technologies to enhance daily workflows and augment problem-solving Distributed Systems
- Experience: Minimum of 7 years of relevant experience in designing, analyzing, and troubleshooting large-scale distributed systems
- Technical Ecosystem Mastery:
- Deep hands-on experience with C, Java, GoLang, Python, Terraform, Ansible, Python automation, networking, Kubernetes, and AWS cloud
- Web Protocols & Security: Comprehensive understanding of web security and core protocols including HTTP, SSL/TLS, DNS, SQL, and networking fundamentals Observability Architecture:
- Proven experience in observability, building complex dashboards, managing Grafana, and maintaining a sharp understanding of SLIs, SLOs, and error budgets
- Modern DevOps Expertise:
- Strong DevOps skills across CI/CD pipelines, Source Control Management (SCM), builds/releases, and Continuous Integration tools/frameworks
- Advanced knowledge of Virtualization, Cloud Architecture, modern Cloud Services, and automated deployment methodologies
- A proven track record of resolving critical escalations and proactively preventing the reoccurrence of incidents through targeted process, monitoring, and reliability improvements
- Active contribution to OS/software packaging and distribution, alongside a passion for mentoring others on SRE best practices within the team
Responsibilities
- We are seeking an experienced Senior Staff, Site Reliability Engineer to join our dynamic SRE Cloud Infrastructure & Operations team.
- In this high-impact role, you will report directly to the Director of Site Reliability Engineering and play a pivotal part in architecting, scaling, and maintaining our next-generation cloud infrastructure.
- You will bridge the gap between development and operations, ensuring our large-scale distributed systems are highly available, secure, and incredibly resilient.
- Architecting & Automating: Design, implement, and manage various advanced cloud management automations to eliminate toil and accelerate delivery
- Container Orchestration: Oversee and optimize containerized architectures using EKS and GKE to ensure robust production performance Observability Systems:
- Lead the creation, deployment, and optimization of highly scalable monitoring and alerting systems
- Cloud Operations & Incident Management:
- Own cloud operations, deployments, on-call support, and incident management while continuously designing and tuning Linux and BSD-based systems
- Cross-Functional Collaboration: Serve as a core member of cross-functional project teams
- contributing to technology-based solutions and consulting on concept feasibility for new initiative
More openings at Zscaler
- Sr. Software Development Engineer (React/Typescript/Browser Extension Development) — Mohali
- Sr. Software Development Engineer - Python Automation / Kubernetes / Networking — Bengaluru
- Sr. Software Development Engineer (Backend - Python/Go) — Mohali
- Sr. Software Development Engineer — India
- Sr. Software Development Engineer — Bengaluru
- Software Development Engineer (DevSecOps) — Bengaluru