Loading open roles
Loading open roles
Loading role

Connect Pro Management Consultants · posted 6 months ago
|
Job Title: |
OpenStack Engineer L3 (SME) |
Band: |
A3/E4 |
|
Location: |
Noida (Onsite) |
Job Type: |
Full Time (FTE) |
|
|
|
|
|
|
Job Summary: |
|||
|
Lead India's Cloud Revolution & Operate at Hyperscale: Join Airtel's New Platform Operations Team!
Airtel is making a strategic leap, launching its own hyperscale cloud service in India. We're building an elite Platform Operations team to be the custodians of this vital infrastructure, delivering unmatched operational excellence from day one.
This role places you at the helm of overseeing a sophisticated and large cloud infrastructure built on cutting-edge technologies. You will be instrumental in creating a world-class core system that guarantee the availability, performance, security, and scalability of our sovereign cloud services, leveraging Airtel's robust network connectivity, extensive infrastructure backbone and deep enterprise relationships.
An L3 OpenStack Engineer is a SME level cloud infrastructure specialist responsible for designing, deploying, managing, and troubleshooting complex OpenStack cloud environments. Their role involves deep expertise in OpenStack components, virtualization technologies, and cloud infrastructure automation, ensuring high availability, performance, and security of private or public clouds.
We are looking for seasoned operations professionals with a passion for running resilient, secure, and highly performance distributed systems. If you are passionate about the challenge of running infrastructure at a gigantic scale while delivering world-class service delivery for a hyperscaler, Airtel Cloud offers a unique opportunity. Come build and run the future with us!
|
|||
|
Key Roles & Responsibilities: |
|||
|
· Design and Implementation o Design and implement scalable, secure, and highly available OpenStack cloud environments. o Configure and manage core OpenStack services such as Nova (compute), Neutron (networking), Cinder (block storage), Keystone (identity), Glance (image), Heat (orchestration), and Swift (object storage). o Integrate OpenStack with third-party tools for monitoring (e.g., Prometheus, Grafana), logging (ELK Stack), and CI/CD pipelines. o Customize OpenStack components to meet specific business needs and workloads o Deploy, configure, and maintain the OpenStack undercloud environment, which is the foundational management layer responsible for deploying and managing the overcloud (production cloud). o Manage and upgrade undercloud services to ensure stable and scalable cloud infrastructure operations. o Automate undercloud deployment processes using tools like TripleO (OpenStack on OpenStack). o Design and deploy overcloud environments tailored to specific workloads, including NFV workloads. o Manage lifecycle operations of the overcloud including scaling, patching, and upgrading o Deploy and manage OVN as the Neutron ML2 mechanism driver to provide logical network abstraction and overlay networking. o Maintain OVN databases (northbound and southbound), OVN controllers, and metadata agents across controller, compute, and gateway nodes. o Implement OVN Layer 3 high availability (L3 HA) features using gateway chassis and BFD monitoring for gateway node health. o Integrate OVN logical networks with physical networks using gateways and tunnels (e.g., Geneve). o Monitor and optimize OVN/Open vSwitch flows and OpenFlow rules for efficient network traffic handling o Understand and configure OpenDaylight controller architecture and components to manage SDN environments. o Integrate OpenDaylight with OpenStack for advanced SDN use cases, including network automation and programmability. o Develop and maintain network scenarios leveraging OpenDaylight for dynamic network control o Deploy and manage OVN as the Neutron ML2 mechanism driver to provide logical network abstraction and overlay networking. · Administration and Maintenance o Perform day-to-day administration and monitoring of OpenStack deployments to ensure optimal performance and availability. o Audit OpenStack environments regularly for vulnerabilities and remediate security issues. o Monitor system health, usage, and performance metrics; set up alerts and dashboards for proactive issue detection. o Maintain and upgrade OpenStack software components and ensure compliance with security best practices · Automation and Optimization o Develop and maintain automation scripts and tools (using Ansible, Terraform, Python, Bash, or Go) for efficient cloud management and deployment automation. o Optimize cloud infrastructure cost, availability, and performance. o Manage infrastructure as code and integrate continuous integration/continuous deployment (CI/CD) pipelines. o Investigate and implement new features or technologies to improve scalability and efficiency · Security, Compliance & Best Practices o Implement and enforce strict security guidelines, including patch management, server permissions, and compliance with ISO/ITSM standards o Develop and execute disaster recovery plans and high-availability strategies for Wintel environments. o Ensure adherence to change management processes and documentation standards. · Troubleshooting & Support o Provide Tier III (L3) production support, handling escalations and resolving complex technical issues. o Collaborate with Site Reliability Engineering (SRE) teams and other engineering groups to ensure smooth integration and operation. o Troubleshoot networking, storage, compute, and security issues within the OpenStack environment. o Validate OpenStack releases through automated functional and performance testing · Collaboration & Continuous Improvement o Collaborate with cross-functional IT teams to integrate automation and virtualization solutions, including OpenStack where applicable. o Participate in technology roadmap planning and provide recommendations for continuous improvement. o Mentor and train junior staff, fostering skills development and knowledge sharing · Research and Innovation o Evaluate and recommend new tools, technologies, and processes to improve OpenStack cloud infrastructure. o Stay updated with latest OpenStack releases, Kubernetes, containerization, and cloud-native technologies. o Support integration with container orchestration platforms like Kubernetes and virtualization technologies such as KVM, VMware
|
|||
|
Technical Skills and Competencies: |
|||
|
· Networking and Storage: Strong understanding of advanced networking concepts (VLANs, routing, firewalls, load balancing) and enterprise storage technologies (SAN - FC/iSCSI, NAS - NFS/CIFS, Object Storage - S3) and their impact on backup performance and design. · Databases and Applications: Experience designing backup strategies for critical enterprise applications and databases (e.g., Oracle, Mongo, SQL Server, SAP HANA, NoSQL). · Security Principles: Strong understanding of security concepts (encryption, key management, access control, vulnerability management) applied to data protection.
Core Competencies: · Robust Foundation in Core Cloud Computing Principles: o Exhibits a strong theoretical and practical understanding of IaaS, PaaS, SaaS, Compute, Storage, Networking, Security, Databases, and Serverless architectures. o Demonstrates a proactive stance towards acquiring knowledge of new cloud services and technologies, actively seeking opportunities to optimize processes, enhance tooling, and improve system reliability. o Experience in cloud platforms (AWS, Azure, GCP) is a plus. · Expertise in Developing Automation Solutions: o Exhibits the ability to design, develop, and implement scripts to streamline repetitive operational tasks, automate deployments, and enhance monitoring capabilities. o Proficiency in at least one scripting language relevant to automation (e.g., Python, Bash, PowerShell, Ruby, and Perl). · Systematic and Data-Driven Technical Diagnostics and Remediation: o Employs a methodical and rigorous approach to identify, diagnose, and resolve technical challenges within distributed systems o Leverage data analysis of logs, performance metrics, and system behavior to pinpoint underlying root causes. · Commitment to Operational Excellence and Process Adherence: o Demonstrates a strong commitment to following established operational procedures, including incident, change, problem, and release management in alignment with ITIL frameworks o Exhibit comprehensive understanding and practical application of the incident lifecycle, encompassing systematic procedures for identification, containment, eradication, recovery, and post-incident analysis. o Proficient in using enterprise-grade monitoring tools (e.g., Prometheus, Grafana, Zabbix, SolarWinds, ELK) and ITSM/ticketing systems (e.g., ServiceNow, Jira) for proactive incident detection, alert management, and streamlined issue resolution workflows. · Exceptional Communicator and Collaborative Team Player: o Demonstrates mastery of English communication, both written and verbal, alongside exceptional interpersonal skills that facilitate strong working relationships o Drive effective collaboration and teamwork across diverse engineering teams to achieve shared goals. · Ability to Thrive in a Dynamic Support Environment: o Exhibits the flexibility and commitment required to work in a 24/7 support model, including rotational shifts and on-call responsibilities. · Experience in Infrastructure and Procedural Documentation: o Proven ability to create and maintain comprehensive documentation for infrastructure, operational procedures, and incident resolutions. |
|||
|
Qualifications: |
|||
|
· Overall 10-14 years of extensive experience and 7+ years OpenStack cloud platforms experience · Bachelor's or master's degree in computer science, Information Technology, or a related field. · Red Hat Certified Specialist in Red Hat OpenStack (RHOSP) or equivalent certification. · Demonstrated experience leading technical teams or projects. |
|||