Senior Site Reliability Engineer
Barclays · Pune · Posted 2026-07-18
Tech stack: Go, Python, REST API
Apply on the company site · Get a referral for this role
Barclays salary & ratings · Barclays interview process · More live openings
About the role
To apply software engineering techniques, automation, and best practices in incident response, to ensure the reliability, availability, and scalability of the systems, platforms, and technology through them.
Responsibilities:
- Availability, performance, and scalability of systems and services through proactive monitoring, maintenance, and capacity planning.
- Resolution, analysis and response to system outages and disruptions, and implement measures to prevent similar incidents from recurring.
- Development of tools and scripts to automate operational processes, reducing manual workload, increasing efficiency, and improving system resilience.
- Monitoring and optimisation of system performance and resource usage, identify and address bottlenecks, and implement best practices for performance tuning.
- Collaboration with development teams to integrate best practices for reliability, scalability, and performance into the software development lifecycle, and work closely with other teams to ensure smooth and efficient operations.
- Stay informed of industry technology trends and innovations, and actively contribute to the organization's technology communities to foster a culture of technical excellence and growth.
Qualifications:
- Should have expertise in languages such as Python, Powershell, or Go, which are essential for automating routine tasks and system deployments.
- Should have ability to manage incidents effectively, troubleshoot issues swiftly, and perform root cause analysis to prevent future incidents.
- Deep understanding of systems engineering, including operating systems, networking, and cloud infrastructure. Proficiency in automation tools is crucial for maintaining system reliability at scale.
- Should be able to communicate effectively with team members and stakeholders, ensuring alignment, inspiring and motivating them to embrace new mindsets, cultures, and SRE working practices. This skill is crucial for driving meaningful change and fostering a collaborative environment where innovative ideas can thrive.
- Should be able to familiar with cloud platforms and services, which is increasingly important as more infrastructure moves to the cloud.
- Should have capability to approach problems methodically and find effective solutions, which is vital for maintaining system reliability.
- Experience working with IaaS and/or PaaS products, including some experience of either virtualization, containerization, orchestration of compute/network/storage.
- Experience working with REST API development and Integration, Database management / Query language.
- Proficiency in implementing monitoring and alerting systems to ensure service availability and performance
- You may be assessed on the key critical skills relevant for success in role, such as risk and controls, change and transformation, business acumen strategic thinking and digital and technology, as well as job-specific technical skills.
- This role is based in our Pune office.
Qualifications
- Should have expertise in languages such as Python, Powershell, or Go, which are essential for automating routine tasks and system deployments.
- Should have ability to manage incidents effectively, troubleshoot issues swiftly, and perform root cause analysis to prevent future incidents.
- Deep understanding of systems engineering, including operating systems, networking, and cloud infrastructure.
- Proficiency in automation tools is crucial for maintaining system reliability at scale.
- Should be able to communicate effectively with team members and stakeholders, ensuring alignment, inspiring and motivating them to embrace new mindsets, cultures, and SRE working practices.
- This skill is crucial for driving meaningful change and fostering a collaborative environment where innovative ideas can thrive.
- Should be able to familiar with cloud platforms and services, which is increasingly important as more infrastructure moves to the cloud.
- Should have capability to approach problems methodically and find effective solutions, which is vital for maintaining system reliability.
- Experience working with IaaS and/or PaaS products, including some experience of either virtualization, containerization, orchestration of compute/network/storage.
- Experience working with REST API development and Integration, Database management / Query language.
- Proficiency in implementing monitoring and alerting systems to ensure service availability and performance
- You may be assessed on the key critical skills relevant for success in role, such as risk and controls, change and transformation, business acumen strategic thinking and digital and technology, as well as job-specific technical skills.
- This role is based in our Pune office.
Responsibilities
- Availability, performance, and scalability of systems and services through proactive monitoring, maintenance, and capacity planning.
- Resolution, analysis and response to system outages and disruptions, and implement measures to prevent similar incidents from recurring.
- Development of tools and scripts to automate operational processes, reducing manual workload, increasing efficiency, and improving system resilience.
- Monitoring and optimisation of system performance and resource usage, identify and address bottlenecks, and implement best practices for performance tuning.
- Collaboration with development teams to integrate best practices for reliability, scalability, and performance into the software development lifecycle, and work closely with other teams to ensure smooth and efficient operations.
- Stay informed of industry technology trends and innovations, and actively contribute to the organization's technology communities to foster a culture of technical excellence and growth.