Infrastructure and Enterprise Engineer III
Expedia · Gurugram · 8+ yrs experience · Posted 2026-07-18
Tech stack: Python
Apply on the company site · Get a referral for this role
Expedia salary & ratings · Expedia interview process · More live openings
About the role
Here, you'll do meaningful work that helps millions of people discover, book, and experience travel with more ease, confidence, and joy. Our five Behaviors-Traveler First, Think Big, Operate with Excellence, Ownership Mindset, and Succeed Together-help foster a supportive environment where people can grow their careers and have the flexibility, benefits, and support to do their best work. Join us and build for travelers everywhere.
Responsibilities:
- Design and build workflow automation that removes toil, improves operational consistency, and accelerates incident response
- Develop AI-enabled operational capabilities that improve detection, diagnosis, communication, and recovery workflows
- Build and maintain services, tools, and automations using Python as a primary engineering language
- Use platforms like n8n and adjacent automation tooling to orchestrate operational workflows across systems
- Strengthen observability across metrics, logs, traces, alerting, and service health signals
- Partner with SRE, platform, infrastructure, and service owners to improve
- resilience, MTTR, and engineering standards
- Identify repeatable failure patterns and turn them into durable automation,
- guardrails, and self-healing opportunities
- Contribute to DevOps practices including CI/CD, release reliability, infrastructure automation, and production readiness
- Raise the technical bar through design reviews, code quality, operational rigor, and mentoring of other engineers.
- Eliminate manual steps in recurring operations processes AI enablement
- Apply AI to operational workflows such as summarization, triage assistance,
- intelligent routing, runbook guidance, anomaly context, and post-incident analysis
Qualifications:
- Software design and hands-on implementation experience
- 8+ years of experience in reliability engineering and automation
- Strong Python development skills, including building automation frameworks,
- services, APIs, and production-grade tooling
- Strong experience with workflow automation, including event-driven and API-led integrations
- Deep understanding of observability, including logs, metrics, traces, alerting
- strategy, and operational telemetry
- Strong DevOps background, including CI/CD, infrastructure automation, production support, and release reliability
- Experience working in high-availability or mission-critical environments
- Strong debugging, systems thinking, and incident management instincts
- Ability to work cross-functionally and influence technical direction.
- Experience with AIOps, operational intelligence, or AI/LLM-enabled engineering workflows
- Familiarity with incident management, problem management, and post-incident review practices
- Experience integrating across tools used in enterprise operations ecosystems
- Knowledge of SRE practices such as SLIs, SLOs, error budgets, and reliability
- engineering tradeoffs
- Experience mentoring engineers and leading medium-to-large technical initiatives
- Exposure to cloud-native infrastructure, platform engineering, and resilience
- automation patterns
Qualifications
- Software design and hands-on implementation experience
- 8+ years of experience in reliability engineering and automation
- Strong Python development skills, including building automation frameworks services, APIs, and production-grade tooling
- Strong experience with workflow automation, including event-driven and API-led integrations
- Deep understanding of observability, including logs, metrics, traces, alerting strategy, and operational telemetry
- Strong DevOps background, including CI/CD, infrastructure automation, production support, and release reliability
- Experience working in high-availability or mission-critical environments
- Strong debugging, systems thinking, and incident management instincts
- Ability to work cross-functionally and influence technical direction.
- Experience with AIOps, operational intelligence, or AI/LLM-enabled engineering workflows
- Familiarity with incident management, problem management, and post-incident review practices
- Experience integrating across tools used in enterprise operations ecosystems
- Knowledge of SRE practices such as SLIs, SLOs, error budgets, and reliability engineering tradeoffs
- Experience mentoring engineers and leading medium-to-large technical initiatives
- Exposure to cloud-native infrastructure, platform engineering, and resilience automation patterns
Responsibilities
- Design and build workflow automation that removes toil, improves operational consistency, and accelerates incident response
- Develop AI-enabled operational capabilities that improve detection, diagnosis, communication, and recovery workflows
- Build and maintain services, tools, and automations using Python as a primary engineering language
- Use platforms like n8n and adjacent automation tooling to orchestrate operational workflows across systems
- Strengthen observability across metrics, logs, traces, alerting, and service health signals
- Partner with SRE, platform, infrastructure, and service owners to improve resilience, MTTR, and engineering standards
- Identify repeatable failure patterns and turn them into durable automation guardrails, and self-healing opportunities
- Contribute to DevOps practices including CI/CD, release reliability, infrastructure automation, and production readiness
- Raise the technical bar through design reviews, code quality, operational rigor, and mentoring of other engineers.
- Eliminate manual steps in recurring operations processes AI enablement
- Apply AI to operational workflows such as summarization, triage assistance intelligent routing, runbook guidance, anomaly context, and post-incident analysis