Principal Site Reliability Engineer (Linux/Networking/Automation)

Zscaler · Hyderabad · 10+ yrs experience · Posted 2026-07-18

Tech stack: AWS, Ansible, Kubernetes, Linux, Python, SQL, Terraform

Apply on the company site · Get a referral for this role

Zscaler salary & ratings · Zscaler interview process · More live openings

About the role

Responsibilities:
- We are looking for a Principal Site Reliability Engineer to join our team.
- This is a hybrid role based in Hyderabad, reporting to the Director, Site Reliability Engineering in the Engineering department.
- You will contribute as a development engineer within our Engineering team, helping to build and enhance the world’s largest cloud security platform.
- You will bring your vision and passion to a team of experts enabling organizations worldwide to harness speed and agility through a cloud-first strategy.
- Perform operational duties for FedRAMP cloud products, including deployments, on-call support, and incident management
- Join deployment sync calls and conduct Operations hand-offs to ensure seamless continuity
- Manage cloud infrastructure elements, including AWS GovCloud, private cloud, containers, and VMs
- Operate and enhance monitoring systems while driving automation, scripting, and Infrastructure as Code (IaC) efforts
- Write and maintain documentation, resolve escalations, prevent incident recurrence, and promote DevOps best practices
Qualifications:
- Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain
- 10+ years of experience as a Site Reliability Engineer with expertise in Operations and Engineering
- Experience with FedRAMP compliance (High/Moderate levels), vulnerability management, and continuous monitoring, including scanning, patching, and reporting
- Proficiency in Linux administration, network troubleshooting, and infrastructure as code (Ansible, Terraform) in cloud environments
- Experience in large-scale distributed systems, containerized architectures (AWS ECS, Kubernetes), and cloud services, with a strong foundation in web security, networking, and coding (Python)
- Experience implementing AIOps frameworks, leveraging machine learning for predictive cloud infrastructure autoscaling, or utilizing AI-driven log anomaly detection tools to optimize root-cause analysis
- Experience with containerized architectures such as AWS ECS and Kubernetes
- Knowledge of web security protocols including HTTP, SSL/TLS, DNS, SQL, and networking fundamentals

Qualifications

- Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain
- 10+ years of experience as a Site Reliability Engineer with expertise in Operations and Engineering
- Experience with FedRAMP compliance (High/Moderate levels), vulnerability management, and continuous monitoring, including scanning, patching, and reporting
- Proficiency in Linux administration, network troubleshooting, and infrastructure as code (Ansible, Terraform) in cloud environments
- Experience in large-scale distributed systems, containerized architectures (AWS ECS, Kubernetes), and cloud services, with a strong foundation in web security, networking, and coding (Python)
- Experience implementing AIOps frameworks, leveraging machine learning for predictive cloud infrastructure autoscaling, or utilizing AI-driven log anomaly detection tools to optimize root-cause analysis
- Experience with containerized architectures such as AWS ECS and Kubernetes
- Knowledge of web security protocols including HTTP, SSL/TLS, DNS, SQL, and networking fundamentals

Responsibilities

- We are looking for a Principal Site Reliability Engineer to join our team.
- This is a hybrid role based in Hyderabad, reporting to the Director, Site Reliability Engineering in the Engineering department.
- You will contribute as a development engineer within our Engineering team, helping to build and enhance the world’s largest cloud security platform.
- You will bring your vision and passion to a team of experts enabling organizations worldwide to harness speed and agility through a cloud-first strategy.
- Perform operational duties for FedRAMP cloud products, including deployments, on-call support, and incident management
- Join deployment sync calls and conduct Operations hand-offs to ensure seamless continuity
- Manage cloud infrastructure elements, including AWS GovCloud, private cloud, containers, and VMs
- Operate and enhance monitoring systems while driving automation, scripting, and Infrastructure as Code (IaC) efforts
- Write and maintain documentation, resolve escalations, prevent incident recurrence, and promote DevOps best practices

More openings at Zscaler