Staff Cloud Backend Engineer- Observability and Site Reliability
Coupang · Bengaluru · 12+ yrs experience · Posted 2026-07-18
Tech stack: AWS, Azure, GCP, Go, Kubernetes, Python, Backend
Apply on the company site · Get a referral for this role
Coupang salary & ratings · Coupang interview process · More live openings
About the role
Responsibilities:
- As a Staff Data Centre Observability and Site Reliability Engineer, you will own the design and operation of scalable observability platforms to ensure the reliability, performance, and availability of datacentre services.
- You will apply SRE best practices, automation, and performance optimization to deliver resilient infrastructure.
- This role partners closely with engineering teams and vendors to drive operational excellence while maintaining security and compliance standards.
- Observability and Monitoring:
- Design, implement, and maintain observability solutions for datacentre infrastructure.
- Develop, deploy, and maintain the operational and reliability components of a large-scale Observability and Telemetry collection platform, emphasizing performance at scale, real-time monitoring, logging, and alerting.
- Participate in and enhance the entire lifecycle of services, from inception and design to deployment, operation, and refinement.
- Develop and optimize monitoring systems to ensure high availability and performance.
- Create and manage dashboards, alerts, and reports to provide visibility into system health and performance.
- Site Reliability Engineering (SRE):
- Implement SRE best practices to improve the reliability, scalability, and performance of datacentre services.
- Develop and maintain automation scripts for infrastructure provisioning, monitoring, and management.
- Conduct root cause analysis and post-mortem reviews to prevent recurrence of incidents.
- Performance Optimization:
- Analyze and optimize the performance of datacentre systems and applications.
- Implement best practices for resource utilization and efficiency.
- Collaboration:
- Work closely with other engineering teams to understand and meet their observability and reliability requirements.
- Collaborate with hardware and software vendors to evaluate and integrate new technologies.
- Security and Compliance:
- Ensure that observability and reliability solutions comply with security policies and industry standards.
- Implement and maintain security measures to protect data and infrastructure.
- Troubleshooting and Support:
- Provide support for observability and reliability-related issues, including debugging and resolving hardware and software problems.
- Develop and maintain documentation for troubleshooting procedures and best practices.
- Continuous Improvement:
- Stay updated with the latest advancements in observability and SRE technologies and integrate them into the infrastructure.
- Continuously improve the reliability, scalability, and performance of datacentre services.
Qualifications:
- Education: Bachelor’s or Master’s degree in Computer Science, Engineering, or a related technical field.
- Experience: 8–12 years of progressive software engineering experience, with a heavy emphasis on distributed systems, cloud-native architectures, or platform operations.
- Programming: Strong proficiency in Go or Python, with a deep understanding of networked systems and performance optimization.
- Orchestration: Expert-level knowledge of Kubernetes internals (scheduling, controllers) and containerization ecosystems.
- Traffic Management: Proven experience with load balancing, service mesh, and request routing at scale.
- Operational Excellence: A strong "ownership" mindset with a track record of maintaining mission-critical, high-availability systems in production.
- AI/ML Domain Knowledge: Prior experience building infrastructure specifically for LLM inference or large-scale training clusters.
- Low-Level Optimization: Familiarity with inference, including mixed precision, kernel tuning, or custom hardware accelerators.
- Public/Private Cloud: Experience managing hybrid-cloud or multi-AZ deployments across AWS, Azure, or GCP.
- Compliance: Experience operating in regulated environments with strict security and compliance requirements.
- Hybrid
Qualifications
- Education: Bachelor’s or Master’s degree in Computer Science, Engineering, or a related technical field.
- Experience: 8–12 years of progressive software engineering experience, with a heavy emphasis on distributed systems, cloud-native architectures, or platform operations.
- Programming: Strong proficiency in Go or Python, with a deep understanding of networked systems and performance optimization.
- Orchestration: Expert-level knowledge of Kubernetes internals (scheduling, controllers) and containerization ecosystems.
- Traffic Management: Proven experience with load balancing, service mesh, and request routing at scale.
- Operational Excellence: A strong "ownership" mindset with a track record of maintaining mission-critical, high-availability systems in production.
- AI/ML Domain Knowledge:
- Prior experience building infrastructure specifically for LLM inference or large-scale training clusters.
- Low-Level Optimization: Familiarity with inference, including mixed precision, kernel tuning, or custom hardware accelerators.
- Public/Private Cloud: Experience managing hybrid-cloud or multi-AZ deployments across AWS, Azure, or GCP.
- Compliance: Experience operating in regulated environments with strict security and compliance requirements.
- Hybrid
Responsibilities
- As a Staff Data Centre Observability and Site Reliability Engineer, you will own the design and operation of scalable observability platforms to ensure the reliability, performance, and availability of datacentre services.
- You will apply SRE best practices, automation, and performance optimization to deliver resilient infrastructure.
- This role partners closely with engineering teams and vendors to drive operational excellence while maintaining security and compliance standards.
- Observability and Monitoring:
- Design, implement, and maintain observability solutions for datacentre infrastructure.
- Develop, deploy, and maintain the operational and reliability components of a large-scale Observability and Telemetry collection platform, emphasizing performance at scale, real-time monitoring, logging, and alerting.
- Participate in and enhance the entire lifecycle of services, from inception and design to deployment, operation, and refinement.
- Develop and optimize monitoring systems to ensure high availability and performance.
- Create and manage dashboards, alerts, and reports to provide visibility into system health and performance.
- Site Reliability Engineering (SRE):
- Implement SRE best practices to improve the reliability, scalability, and performance of datacentre services.
- Develop and maintain automation scripts for infrastructure provisioning, monitoring, and management.
- Conduct root cause analysis and post-mortem reviews to prevent recurrence of incidents.
- Performance Optimization: Analyze and optimize the performance of datacentre systems and applications.
- Implement best practices for resource utilization and efficiency.
- Collaboration: Work closely with other engineering teams to understand and meet their observability and reliability requirements.
- Collaborate with hardware and software vendors to evaluate and integrate new technologies.
- Security and Compliance:
- Ensure that observability and reliability solutions comply with security policies and industry standards.
- Implement and maintain security measures to protect data and infrastructure.
- Troubleshooting and Support:
- Provide support for observability and reliability-related issues, including debugging and resolving hardware and software problems.
- Develop and maintain documentation for troubleshooting procedures and best practices.
- Continuous Improvement: Stay updated with the latest advancements in observability and SRE technologies and integrate them into the infrastructure.
- Continuously improve the reliability, scalability, and performance of datacentre services.