Senior Site Reliability Engineer

Akamai · India · 5+ yrs experience · Posted 2026-07-18

Tech stack: JavaScript, Linux, Python, SQL, Unix

Apply on the company site · Get a referral for this role

Akamai salary & ratings · Akamai interview process · More live openings

About the role

Do you like collaborating across teams to solve complex problems?
Do you enjoy solving large scale distributed content delivery challenges?
Join our critical Platform and Reliability Engineering Team!
The Platform & Reliability Engineering team is responsible for defining, measuring, & optimizing the key performance indicators of delivery customers. Your expertise in software engineering and systems administration will be instrumental in building robust and resilient infrastructure.
Responsibilities:
- Enhance system reliability, scalability, and performance by designing, implementing, and maintaining robust infrastructure and operational processes.
- Collaborate with cross-functional teams to ensure seamless deployment, monitoring, and incident response. Proactively identify and resolve system issues to support business continuity and growth. Advocate for automation and best practices to optimize system efficiency, availability, and security across the organization.
- Working on Internet technologies to improve the performance, availability, and scalability of large distributed content delivery systems.
- Collaborating with cross-functional teams, including Product and engineering, to define and establish measurable Service Level Indicators and Objectives.
- Providing technical expertise and feedback to ensure system designs and implementations align with reliability and performance requirements effectively.
- Monitoring platform availability and performance, analyzing data to debug issues, and implementing corrective actions to prevent future occurrences.
- Developing and implementing automation solutions aimed at enhancing operational efficiency while minimizing repetitive tasks.
- Participating in design reviews and providing technical guidance to ensure designs meet requirements for scalability, performance, and robustness
- Staying updated on recent developments in cloud computing, DevOps, and SRE best practices without missing emerging trends.
Qualifications:
- Have 5+ years of relevant experience and a Bachelor's degree in Computer Science, Engineering, or related field.
- Demonstrate expertise in scripting languages like Python, Bash, and JavaScript to enable automation and create efficient tools.
- Utilize monitoring and alerting tools such as Prometheus, Grafana, ADBMS, and Datadog effectively for creating and managing dashboards.
- Work proficiently in UNIX/Linux environments, showcasing exceptional problem-solving abilities and comprehensive system expertise.
- Utilize Oracle SQL to perform data integrity checks, identify root causes of anomalies, and generate detailed reports.
- Improve outcomes, embrace ongoing learning, and deliver automation-centered operational excellence across various initiatives with proactive self-direction.
- Demonstrate customer focus, accountability, and exceptional communication and teamwork abilities across various cross-functional groups.

Qualifications

- Have 5+ years of relevant experience and a Bachelor's degree in Computer Science, Engineering, or related field.
- Demonstrate expertise in scripting languages like Python, Bash, and JavaScript to enable automation and create efficient tools.
- Utilize monitoring and alerting tools such as Prometheus, Grafana, ADBMS, and Datadog effectively for creating and managing dashboards.
- Work proficiently in UNIX/Linux environments, showcasing exceptional problem-solving abilities and comprehensive system expertise.
- Utilize Oracle SQL to perform data integrity checks, identify root causes of anomalies, and generate detailed reports.
- Improve outcomes, embrace ongoing learning, and deliver automation-centered operational excellence across various initiatives with proactive self-direction.
- Demonstrate customer focus, accountability, and exceptional communication and teamwork abilities across various cross-functional groups.

Responsibilities

- Enhance system reliability, scalability, and performance by designing, implementing, and maintaining robust infrastructure and operational processes.
- Collaborate with cross-functional teams to ensure seamless deployment, monitoring, and incident response.
- Proactively identify and resolve system issues to support business continuity and growth.
- Advocate for automation and best practices to optimize system efficiency, availability, and security across the organization.
- Working on Internet technologies to improve the performance, availability, and scalability of large distributed content delivery systems.
- Collaborating with cross-functional teams, including Product and engineering, to define and establish measurable Service Level Indicators and Objectives.
- Providing technical expertise and feedback to ensure system designs and implementations align with reliability and performance requirements effectively.
- Monitoring platform availability and performance, analyzing data to debug issues, and implementing corrective actions to prevent future occurrences.
- Developing and implementing automation solutions aimed at enhancing operational efficiency while minimizing repetitive tasks.
- Participating in design reviews and providing technical guidance to ensure designs meet requirements for scalability, performance, and robustness
- Staying updated on recent developments in cloud computing, DevOps, and SRE best practices without missing emerging trends.

More openings at Akamai