Job Description
Head – Fault Management and Service Assurance (AVP/VP)
Head – Fault Management and Service Assurance is accountable for end-to-end fault lifecycle management of Airtel’s transport networks, ensuring high availability, low latency, rapid restoration, and minimal service impact across Mobility, Broadband, Enterprise, and Homes services.
The role leads Airtel’s transition from reactive transport fault handling to predictive, AI-driven, and autonomous transport operations, aligned with TM Forum Open Digital Architecture (ODA) and accordingly shall drive Autonomous Network maturity targets.
Key Responsibilities
1. End-to-End Transport Fault Management Ownership
- Own and govern L2–L3 Fault Management processes for transport network and all underlying infrastructure platforms, ensuring uniform and standardized execution across regions and vendors.
- Define, own, and continuously improve the end-to-end fault management lifecycle (detect, correlate, diagnose, restore, close) in alignment with TM Forum eTOM Service Assurance and Fault Management components.
- Exercise command ownership of critical infrastructure and service outages (fiber cuts, router failures, power issues, etc.).
- Drive service-impact-based prioritization across Mobile, FTTH, Enterprise, and Home services.
- Lead war rooms during major incidents, co-ordinating with internal teams, Ericsson GNOC, vendors
- Manage escalations, and executive-level updates.
- Ensure SLA / OLA adherence for internal teams and partners.
Establish and govern standardized fault management:
- Alarm & Event Management (Optical, IP/MPLS, etc, cloud infrastructure and application monitoring etc)
- Root Cause Analysis (RCA) & Problem Management
- Proactive fault prevention & service resilience planning across hybrid environments
2. Multi-Domain Coverage with focus on transport
Provide unified fault oversight and drive cross-layer correlation (Physical → Optical → IP → Service(including RAN and Core) across entire technology stack
- Physical infrastructure
- IP/MPLS Core & Aggregation
- DWDM / OTN optical networks
- National Long-Haul, Backhaul, and Submarine cable systems
- RAN and Core networks
- Application and service layers
3. Incident, Problem, and Change Integration
- Ensure seamless integration between fault, incident, problem, and change management processes, including automated ticket creation, routing, and closure.
- Lead or sponsor major incident bridges for high-severity outages, coordinating engineering, field operations, vendors, and partners.
- Drive Problem Management for recurring and chronic issues through structured RCAs, corrective and preventive actions, and closure tracking.
4. Alarm and Event Management
- Govern the complete alarm and event lifecycle, including creation, enrichment, prioritization, correlation, deduplication, clearance, and archival.
- Drive alarm quality improvements and noise reduction using TMF642-aligned severity models and event semantics.
5. AI-Driven & Autonomous Transport Operations
Lead adoption of AIOps for Unified service and infrastructure Management and align transport assurance to Autonomous Network Levels (AN L2–L4+):
- Intelligent alarm correlation & noise suppression
- Fiber cut localization & blast radius analysis
- Predictive service degradation detection (BER, latency, packet loss)
- Closed-loop rerouting, protection switching, and auto-restoration, self-healing and auto scaling capabilities
6. Tooling, Architecture & ODA Alignment
Own the Unified service and infrastructure assurance tooling strategy, including:
- Optical NMS / EMS platforms
- IP/MPLS NMS & telemetry platforms
- Centralized Event Management & ITSM integration
Ensure:
- API-driven, cloud-native, decoupled architecture
- Adoption of TMF Open APIs (Alarm, Event, Trouble Ticket, Performance)
- Alignment with TMF ODA components and reduction of tool silos
7. KPI Ownership & Operational Excellence
Define, track, and improve KPIs, and establish data-driven NOC performance management:
- MTTR / MTTI / MTTA
- Network & Ring Availability
- Fiber Cut Frequency & Restoration Time
- Repeat Transport Incidents
- Customer Impact Minutes
- Alarm noise
8. Vendor & Partner Governance
- Manage OEMs.
- Enforce SLAs, penalties, and service credits.
- Drive elimination of chronic faults through structural and permanent fixes.
- Align vendor roadmaps with Airtel’s transport network evolution.
9. Post-Incident Learning & Continuous Improvement
- Institutionalize comprehensive RCAs for all critical incidents
- Maintain a Problem Backlog for all recurring infrastructure and service issues
- Identify automation and resilience improvements from incidents.
- Publish executive-level service stability insights, providing a holistic view of performance across all domains
10. Team Management and Capability Building
- Build, develop, and retain high-performing NOC fault management and service assurance teams with clearly defined L1/L2/L3 roles and competency frameworks.
- Drive continuous upskilling across network technologies, OSS tools, analytics platforms, and industry frameworks (TM Forum, ITIL).
- Build expertise in optical, IP, telemetry, and AI practices.
- Foster a culture of ownership, speed, and zero-defect mindset.
Key Stakeholders
- Transport Planning & Engineering
- Mobility, FTTH & Enterprise Operations
- IT / Digital Platforms
- Circle & Regional Operations Teams
- Vendors, Fiber & Submarine Partners
- CX & Service Assurance Teams
Skills & Competencies
Core Skills
- Strong understanding of telecom network architectures and fault behaviors across mobile, fixed, IP, transport, and cloud domains.
- Expertise in alarm and fault management platforms, correlation engines, monitoring tools, and ticketing systems.
- Proven leadership in crisis management, stakeholder communication, analytical problem-solving, and customer-centric operations.
Technical
- Deep expertise in Transport Networks (IP/MPLS, Optical)
- Strong knowledge of fault, incident, and problem management
- Experience with AIOps and transport analytics
- Understanding of ITIL, SRE, and observability concepts
- Familiarity with TM Forum frameworks and standards
Leadership
- Crisis and outage management
- Executive communication
- Large-scale operations leadership
- Transformation and change management
Experience & Qualifications
- 15+ years of experience in Telecom Transport Operations
- 5+ years leading large fault management or service assurance teams
- Experience operating Tier-1 national transport networks
- TM Forum / ITIL certifications preferred