Senior Staff Cloud Backend Engineer - Observability and Site Reliability
Coupang · Bengaluru · 12+ yrs experience · Posted 2026-07-18
Tech stack: AWS, Azure, Docker, GCP, Go, Kubernetes, Python, Terraform, Backend
Apply on the company site · Get a referral for this role
Coupang salary & ratings · Coupang interview process · More live openings
About the role
Responsibilities:
- As a Senior Staff Data Centre Observability and Site Reliability Engineer, you will design, build, and operate scalable observability and reliability solutions for large-scale datacenter infrastructure.
- This role focuses on developing high-performance monitoring and telemetry platforms, ensuring system reliability, and driving operational excellence through automation, performance optimization, and SRE best practices.
- The ideal candidate will work across the full service lifecycle—design, deployment, and continuous improvement—while collaborating with cross-functional teams to enhance visibility, resilience, and efficiency of critical systems.
- Observability and Monitoring
- Design, implement, and maintain observability solutions for datacenter infrastructure, including monitoring, logging, alerting, and telemetry systems.
- Develop, deploy, and operate large-scale observability and telemetry platforms with a focus on real-time monitoring, high performance, and scalability.
- Own and contribute to the full lifecycle of observability services—from design and development to deployment and ongoing optimization.
- Build and enhance monitoring systems to ensure high availability, reliability, and performance of infrastructure.
- Create and manage dashboards, alerts, and reports to provide clear visibility into system health, performance, and capacity trends.
- Site Reliability Engineering (SRE)
- Apply SRE principles and best practices to improve reliability, scalability, and operational efficiency of datacenter services.
- Develop and maintain automation for infrastructure provisioning, monitoring, and system management.
- Lead root cause analysis (RCA) and post-incident reviews, driving corrective actions to prevent recurrence and improve system resilience.
- Performance Optimization
- Analyze system and application performance across the datacenter infrastructure to identify bottlenecks and improvement areas.
- Implement optimization strategies to enhance performance, efficiency, and resource utilization.
- Collaboration
- Partner with cross-functional engineering teams to understand observability and reliability requirements and deliver effective solutions.
- Collaborate with hardware and software vendors to evaluate, integrate, and optimize new technologies within the ecosystem.
- Security and Compliance
- Ensure observability and reliability solutions adhere to organizational security policies and industry standards.
- Implement and maintain appropriate security controls to safeguard infrastructure, systems, and data.
- Troubleshooting and Support
- Provide hands-on support for observability and reliability issues, including debugging complex hardware and software problems.
- Develop and maintain documentation, including troubleshooting guides and operational best practices, to support efficient issue resolution.
- Continuous Improvement
- Stay current with emerging trends, tools, and technologies in observability and SRE, and incorporate them into the platform.
- Continuously enhance the scalability, reliability, and operational efficiency of datacenter services through proactive improvements.
Qualifications:
- Bachelor’s or Master’s degree in Computer Science, Engineering, or a related technical field.
- 12+ years of progressive software engineering experience, with a heavy emphasis on distributed systems, cloud-native architectures, or platform operations.
- Proven experience in managing and optimizing large-scale datacenter environments
- Strong proficiency in Go or Python, with a deep understanding of networked systems and performance optimization.
- Expert-level knowledge of Kubernetes internals (scheduling, controllers) and containerization ecosystems.
- Proven experience with load balancing, service mesh, and request routing at scale.
- Proficiency in observability tools and technologies (e.g., Prometheus, Grafana, ELK Stack).
- Experience with SRE practices and tools (e.g., Kubernetes, Docker, Terraform).
- Familiarity with cloud platforms (AWS, Azure, GCP) and their observability and reliability services
- Prior experience building infrastructure specifically for LLM inference or large-scale training clusters.
- Familiarity with inference, including mixed precision, kernel tuning, or custom hardware accelerators.
- Experience managing hybrid-cloud or multi-AZ deployments across AWS, Azure, or GCP.
- Experience operating in regulated environments with strict security and compliance requirements
Qualifications
- Bachelor’s or Master’s degree in Computer Science, Engineering, or a related technical field.
- 12+ years of progressive software engineering experience, with a heavy emphasis on distributed systems, cloud-native architectures, or platform operations.
- Proven experience in managing and optimizing large-scale datacenter environments
- Strong proficiency in Go or Python, with a deep understanding of networked systems and performance optimization.
- Expert-level knowledge of Kubernetes internals (scheduling, controllers) and containerization ecosystems.
- Proven experience with load balancing, service mesh, and request routing at scale.
- Proficiency in observability tools and technologies (e.g., Prometheus, Grafana, ELK Stack).
- Experience with SRE practices and tools (e.g., Kubernetes, Docker, Terraform).
- Familiarity with cloud platforms (AWS, Azure, GCP) and their observability and reliability services
- Prior experience building infrastructure specifically for LLM inference or large-scale training clusters.
- Familiarity with inference, including mixed precision, kernel tuning, or custom hardware accelerators.
- Experience managing hybrid-cloud or multi-AZ deployments across AWS, Azure, or GCP.
- Experience operating in regulated environments with strict security and compliance requirements
Responsibilities
- As a Senior Staff Data Centre Observability and Site Reliability Engineer, you will design, build, and operate scalable observability and reliability solutions for large-scale datacenter infrastructure.
- This role focuses on developing high-performance monitoring and telemetry platforms, ensuring system reliability, and driving operational excellence through automation, performance optimization, and SRE best practices.
- The ideal candidate will work across the full service lifecycle—design, deployment, and continuous improvement—while collaborating with cross-functional teams to enhance visibility, resilience, and efficiency of critical systems.
- Observability and Monitoring
- Design, implement, and maintain observability solutions for datacenter infrastructure, including monitoring, logging, alerting, and telemetry systems.
- Develop, deploy, and operate large-scale observability and telemetry platforms with a focus on real-time monitoring, high performance, and scalability.
- Own and contribute to the full lifecycle of observability services—from design and development to deployment and ongoing optimization.
- Build and enhance monitoring systems to ensure high availability, reliability, and performance of infrastructure.
- Create and manage dashboards, alerts, and reports to provide clear visibility into system health, performance, and capacity trends.
- Site Reliability Engineering (SRE)
- Apply SRE principles and best practices to improve reliability, scalability, and operational efficiency of datacenter services.
- Develop and maintain automation for infrastructure provisioning, monitoring, and system management.
- Lead root cause analysis (RCA) and post-incident reviews
- driving corrective actions to prevent recurrence and improve system resilience.
- Performance Optimization Analyze system and application performance across the datacenter infrastructure to identify bottlenecks and improvement areas.
- Implement optimization strategies to enhance performance, efficiency, and resource utilization.
- Collaboration Partner with cross-functional engineering teams to understand observability and reliability requirements and deliver effective solutions.
- Collaborate with hardware and software vendors to evaluate, integrate, and optimize new technologies within the ecosystem.
- Security and Compliance
- Ensure observability and reliability solutions adhere to organizational security policies and industry standards.
- Implement and maintain appropriate security controls to safeguard infrastructure, systems, and data.
- Troubleshooting and Support
- Provide hands-on support for observability and reliability issues, including debugging complex hardware and software problems.
- Develop and maintain documentation, including troubleshooting guides and operational best practices, to support efficient issue resolution.
- Continuous Improvement Stay current with emerging trends, tools, and technologies in observability and SRE, and incorporate them into the platform.
- Continuously enhance the scalability, reliability, and operational efficiency of datacenter services through proactive improvements.
More openings at Coupang
- Senior Staff Backend Engineer (Service Mesh) — India
- Senior Staff Backend Engineer — Bengaluru
- Senior Staff Backend Engineer — Bengaluru
- Senior Staff Backend Engineer — Hyderabad
- Senior Software Engineering Manager, Rocket Growth - Seller Success — Bengaluru
- Senior Software Development Manager - Global Operations Team — Hyderabad