Director - Backend Engineering - AI Infra
Coupang · Bengaluru · 15+ yrs experience · Posted 2026-07-18
Tech stack: AWS, Ansible, Azure, C++, GCP, Go, Kubernetes, Linux, Python, Terraform, Backend
Apply on the company site · Get a referral for this role
Coupang salary & ratings · Coupang interview process · More live openings
About the role
Responsibilities:
- Director of Backend Engineering (AI Infrastructure)
- We are seeking a visionary Director of Backend Engineering to lead the teams responsible for the software "brain" that manages our global AI Physical Infrastructure.
- You will oversee the development of the SDN orchestrators, automated fleet management systems, and the high-performance storage backends that power our AI training and inference clusters.
- Your mission is to abstract the complexity of specialized hardware (NVIDIA/HPC) into a seamless, automated, and hyper-reliable cloud platform.
- Strategic Leadership & Fleet Orchestration
- Software-Defined Infrastructure: Lead the design and delivery of an SDN Orchestrator to automate complex GPU networking (InfiniBand/RoCE/NVLink) and core DC routing.
- Fleet Health Automation: Oversee the development of backend services for GPU Health & Fault Detection, automating the lifecycle from burn-in and diagnostics to global RMA workflows.
- Capacity & Traffic Engineering: Drive the backend logic for global traffic routing, load balancing (NGINX/Kong), and IPAM to ensure zero-bottleneck training environments.
- Data & Storage Systems
- HPC Data Pipelines: Collaborate with storage engineers to build backend interfaces for Parallel File Systems (Lustre, Weka, VAST etc.), ensuring high-throughput data delivery to compute nodes.
- Storage Durability: Direct the backend strategy for AI Object Storage, focusing on high durability and low-latency retrieval for massive datasets.
- Engineering Excellence
- Scalable Architecture: Act as the final technical authority for AI Infra Architecture, ensuring systems are resilient, multi-region, and capable of sub-millisecond coordination.
- DevOps & IaC Culture: Champion a "Hardware-as-Code" mindset, utilizing Python, Ansible, and Terraform to eliminate manual intervention in DC operations.
- Team Development
- Lead a multi-disciplinary org including Backend Developers, SDN Engineers, and Infra Ops teams, AI Infra Engineering
- Establish 24/7 L1/L2/L3 operational standards to maintain > 99.99% availability of the AI fleet.
- Required Qualifications
- Experience: 15+ years in Backend Engineering, with at least 5 years in a leadership role managing complex infrastructure (Cloud, FinTech, or HPC).
- Deep Infrastructure Knowledge: Proven experience with Linux internals, hardware-software interfaces (drivers/firmware), and distributed systems.
- Networking Mastery: Solid understanding of L2/L3 networking, and ideally, specialized fabrics like InfiniBand or RoCE.
- The Stack: Professional proficiency in Python, Go, or C++, and deep experience with Terraform, Kubernetes, and Ansible.
- Large-Scale Data: Experience managing high-performance storage backends (GPFS, Lustre, or equivalent parallel systems).
- Hardware Savvy: You don't just write code; you understand power envelopes, liquid cooling constraints, and GPU architecture (NVIDIA/HPE/Dell).
- Preferred Skills
- Experience building custom SDN controllers or orchestration layers from scratch.
- Direct experience with NVIDIA or GPUDirect technologies.
- Previous success in a "Hyper-scale" environment (AWS, Azure, GCP, Meta, AI Cloouds etc.).
- Application Review - Phone Interview - Onsite (or Virtual Onsite) Interview – Offer
- The exact nature of the recruitment process may vary according to the specific job and may be changed due to scheduling or other circumstances.
- Interview schedules and the results will be informed to the applicant via the e-mail address submitted at the application stage.
- This job posting may be closed prior to the stated end date for application if all openings are filled.
- Coupang has the right to rescind an offer of employment if a candidate is found to have submitted false information as part of the application process.
- https://privacy.coupang.com/en/land/jobs/
Qualifications:
- Experience: 15+ years in Backend Engineering, with at least 5 years in a leadership role managing complex infrastructure (Cloud, FinTech, or HPC).
- Deep Infrastructure Knowledge: Proven experience with Linux internals, hardware-software interfaces (drivers/firmware), and distributed systems.
- Networking Mastery: Solid understanding of L2/L3 networking, and ideally, specialized fabrics like InfiniBand or RoCE.
- The Stack: Professional proficiency in Python, Go, or C++, and deep experience with Terraform, Kubernetes, and Ansible.
- Large-Scale Data: Experience managing high-performance storage backends (GPFS, Lustre, or equivalent parallel systems).
- Hardware Savvy: You don't just write code; you understand power envelopes, liquid cooling constraints, and GPU architecture (NVIDIA/HPE/Dell).
- Preferred Skills
- Experience building custom SDN controllers or orchestration layers from scratch.
- Direct experience with NVIDIA or GPUDirect technologies.
- Previous success in a "Hyper-scale" environment (AWS, Azure, GCP, Meta, AI Cloouds etc.).
- Application Review - Phone Interview - Onsite (or Virtual Onsite) Interview – Offer
- The exact nature of the recruitment process may vary according to the specific job and may be changed due to scheduling or other circumstances.
- Interview schedules and the results will be informed to the applicant via the e-mail address submitted at the application stage.
- This job posting may be closed prior to the stated end date for application if all openings are filled.
- Coupang has the right to rescind an offer of employment if a candidate is found to have submitted false information as part of the application process.
- https://privacy.coupang.com/en/land/jobs/
Qualifications
- Experience: 15+ years in Backend Engineering, with at least 5 years in a leadership role managing complex infrastructure (Cloud, FinTech, or HPC).
- Deep Infrastructure Knowledge:
- Proven experience with Linux internals, hardware-software interfaces (drivers/firmware), and distributed systems.
- Networking Mastery: Solid understanding of L2/L3 networking, and ideally, specialized fabrics like InfiniBand or RoCE.
- The Stack: Professional proficiency in Python, Go, or C++, and deep experience with Terraform, Kubernetes, and Ansible.
- Large-Scale Data: Experience managing high-performance storage backends (GPFS, Lustre, or equivalent parallel systems).
- Hardware Savvy: You don't just write code; you understand power envelopes, liquid cooling constraints, and GPU architecture (NVIDIA/HPE/Dell).
- Experience building custom SDN controllers or orchestration layers from scratch.
- Direct experience with NVIDIA or GPUDirect technologies.
- Previous success in a "Hyper-scale" environment (AWS, Azure, GCP, Meta, AI Cloouds etc.).
- Application Review Phone Interview
- Onsite (or Virtual Onsite) Interview Offer
- The exact nature of the recruitment process may vary according to the specific job and may be changed due to scheduling or other circumstances.
- Interview schedules and the results will be informed to the applicant via the e-mail address submitted at the application stage.
- This job posting may be closed prior to the stated end date for application if all openings are filled.
- Coupang has the right to rescind an offer of employment if a candidate is found to have submitted false information as part of the application process.
- https://privacy.coupang.com/en/land/jobs/
Responsibilities
- Director of Backend Engineering (AI Infrastructure)
- We are seeking a visionary Director of Backend Engineering to lead the teams responsible for the software "brain" that manages our global AI Physical Infrastructure.
- You will oversee the development of the SDN orchestrators, automated fleet management systems, and the high-performance storage backends that power our AI training and inference clusters.
- Your mission is to abstract the complexity of specialized hardware (NVIDIA/HPC) into a seamless, automated, and hyper-reliable cloud platform.
- Strategic Leadership & Fleet Orchestration Software-Defined Infrastructure:
- Lead the design and delivery of an SDN Orchestrator to automate complex GPU networking (InfiniBand/RoCE/NVLink) and core DC routing.
- Fleet Health Automation: Oversee the development of backend services for GPU Health & Fault Detection, automating the lifecycle from burn-in and diagnostics to global RMA workflows.
- Capacity & Traffic Engineering:
- Drive the backend logic for global traffic routing, load balancing (NGINX/Kong), and IPAM to ensure zero-bottleneck training environments.
- Data & Storage Systems
- HPC Data Pipelines:
- Collaborate with storage engineers to build backend interfaces for Parallel File Systems (Lustre, Weka, VAST etc.), ensuring high-throughput data delivery to compute nodes.
- Storage Durability: Direct the backend strategy for AI Object Storage, focusing on high durability and low-latency retrieval for massive datasets.
- Engineering Excellence Scalable Architecture: Act as the final technical authority for AI Infra Architecture, ensuring systems are resilient, multi-region, and capable of sub-millisecond coordination.
- DevOps & IaC Culture: Champion a "Hardware-as-Code" mindset, utilizing Python, Ansible, and Terraform to eliminate manual intervention in DC operations.
- Team Development Lead a multi-disciplinary org including Backend Developers, SDN Engineers, and Infra Ops teams, AI Infra Engineering
- Establish 24/7 L1/L2/L3 operational standards to maintain
- 99.99% availability of the AI fleet.
- Experience: 15+ years in Backend Engineering, with at least 5 years in a leadership role managing complex infrastructure (Cloud, FinTech, or HPC).
- Deep Infrastructure Knowledge:
- Proven experience with Linux internals, hardware-software interfaces (drivers/firmware), and distributed systems.
- Networking Mastery: Solid understanding of L2/L3 networking, and ideally, specialized fabrics like InfiniBand or RoCE.
- The Stack: Professional proficiency in Python, Go, or C++, and deep experience with Terraform, Kubernetes, and Ansible.
- Large-Scale Data: Experience managing high-performance storage backends (GPFS, Lustre, or equivalent parallel systems).
- Hardware Savvy: You don't just write code; you understand power envelopes, liquid cooling constraints, and GPU architecture (NVIDIA/HPE/Dell).
- Experience building custom SDN controllers or orchestration layers from scratch.
- Direct experience with NVIDIA or GPUDirect technologies.
- Previous success in a "Hyper-scale" environment (AWS, Azure, GCP, Meta, AI Cloouds etc.).
- Application Review Phone Interview
- Onsite (or Virtual Onsite) Interview Offer
- The exact nature of the recruitment process may vary according to the specific job and may be changed due to scheduling or other circumstances.
- Interview schedules and the results will be informed to the applicant via the e-mail address submitted at the application stage.
- This job posting may be closed prior to the stated end date for application if all openings are filled.
- Coupang has the right to rescind an offer of employment if a candidate is found to have submitted false information as part of the application process.
- https://privacy.coupang.com/en/land/jobs/
More openings at Coupang
- Director, Back-End Engineering — Hyderabad
- Staff Software Engineer — Bengaluru
- Staff Engineer - Backend (Post Purchase Experience) — Bengaluru
- Staff Engineer - Backend (Finance Platform) — Bengaluru
- Staff Engineer - Backend (Finance Platform) — Bengaluru
- Staff Engineer - Backend (E-Commerce Engineering) — Bengaluru