Loading open roles
Loading open roles
Loading role

Vrinda Global · posted 1 month ago
HDFC Bank is India's largest private sector bank, offering a comprehensive range of financial products and services to our customer base of over 92 million. Our extensive distribution network of 8,919 branches and 21,031 ATMs across 3,836 cities and towns as of August 2024, reaches every corner of the country, making us accessible to millions.
Website - https://www.hdfc.bank.in/
Industry - Banking
Company size - 2,11,178+ employees
Job Title:
Senior Site Reliability Engineer
Job Details:
Business Unit: Tech & Digital
Team: DTIT – Enterprise Factory
Reports to: Lead Site Reliability Engineer
Role Type: Individual Contributor
Location - Guwahati, Assam
*Kindly give consent on mail that you are okay with relocating to the job location i.e., Guwahati, Assam*
Job Purpose:
Analysing, troubleshooting, and designing vital services, platforms, and infrastructure while always thinking about reliability, scalability, resilience, security, and performance.
Job Responsibilities:
Help build a Site Reliability Engineering culture by sharing best practices, approaches, documentation, and code with other engineering teams.
Apply automation and software to any tasks or parts of the system which are performed manually.
Able to troubleshoot complicated, cross-platform issues handling OS, Networking, Database in a cloud-based SaaS environment and handle live production incidents.
Monitor application performance, take steps to improve overall application performance and stability, and follow through with implementation.
Conduct system analysis, configuration management, and develop improvements for system software performance, availability, and reliability.
Design, write, ship, and motivate the creation of software and systems to increase observability, product reliability, and organizational efficiency.
Maintain and monitor deployment, orchestration, of servers, docker containers, databases, and general backend infrastructure.
Develop Run Books/Standard Operating Procedure for recurring Production issues, also working on a permanent solve.Perform Incident Analysis on a regular basis with the intention of preventing and finding a long-term solve for Incidents. Define KPI in the form of RPO/RTO/SLI/SLO/Error Budget.
Apply automation and software to any manually performed tasks or system parts.
Develop Run Books/Standard Operating Procedure for recurring Production issues and work on permanent solutions.
Perform Incident Analysis regularly to prevent and find long-term solutions for Incidents.
Educational Qualifications:
B Tech in Computer Science or related discipline preferred.
Key Skills:
Experience in monitoring and analyzing infrastructure performance using standard performance monitoring tools.
Demonstrable experience in Containerization-Docker and orchestration (Kubernetes).
Experience with Infrastructure As Code (Terraform, Cloud Formation, Ansible).
Knowledge and proven hands-on experience in large-scale databases and distributed technologies, such as Kafka and Confluent Platform Kafka.
Basic programming and scripting skills.
Experience Required:
Total Yrs of experience: 8-13
Major Stakeholders:
Internal:
Product Manager from Digital Factory
Business Analyst from BTG team
Incident Management team
Development Team
Skills
Thanks & Regards ,
Sakshi
Senior HR Recruiter
Ph - 8307315248 | E - sakshi.mishra@vrindaglobal.com
LinkedIn - https://www.linkedin.com/in/sakshi-mishra-563024237/
Vrinda Global Consultants
https://vrindaconsultants.com/