Loading open roles
Loading open roles
Loading role

Leading Investment Bank · posted 4 months ago
Job Summary
We are seeking a highly skilled DevOps / Site Reliability Engineer (SRE) with strong expertise in monitoring, logging, tracing, and observability platforms . The ideal candidate will design, implement, and operate scalable observability and monitoring solutions to ensure high availability, performance, and reliability of business-critical systems.
This role requires hands-on experience with modern observability stacks including Prometheus, Grafana, ELK, OpenTelemetry , and related tools, along with strong operational and automation skills.
Experience – SSE role
4-8 years in DevOps / SRE / Observability roles
Key Responsibilities
DevOps & SRE Responsibilities
· Design, implement, and maintain highly available, scalable, and reliable systems
· Apply SRE principles such as SLIs, SLOs, error budgets, and incident management
· Automate infrastructure provisioning, monitoring, and alerting processes
· Participate in on-call rotations and lead incident triage and root cause analysis (RCA)
· Improve system reliability through proactive monitoring and performance tuning
Monitoring & Observability
· Build and manage end-to-end observability platforms covering:
o Metrics
o Logs
o Traces
· Design and implement monitoring frameworks for cloud-native and distributed systems
· Define alerting strategies to reduce noise and improve mean time to detect (MTTD) and recover (MTTR)
· Create and maintain dashboards, alerts, and runbooks for operational teams
Tooling & Platform Management
· Implement and manage observability tools including:
o Metrics : Prometheus, Thanos
o Visualization : Grafana
o Logging : Elasticsearch, Logstash, Kibana (ELK), Fluentd
o Tracing : Jaeger, Tempo
o Telemetry : OpenTelemetry
· Integrate observability tools with CI/CD, cloud platforms, and Kubernetes environments
· Optimize storage, retention, and performance of observability data
Required Technical Skills
Core Skills
· Strong experience in DevOps and/or SRE roles
· Hands-on expertise in:
o Grafana (dashboards, alerts, data sources)
o ELK Stack (Elasticsearch, Logstash, Kibana)
o Prometheus and metrics-based monitoring
· Proven experience setting up monitoring and observability frameworks from scratch Observability Stack
· Practical knowledge of:
o Prometheus & Thanos
o Grafana
o Loki
o Jaeger & Tempo
o Elasticsearch & Kibana
o Fluentd
o OpenTelemetry (instrumentation and exporters)
· Understanding of distributed tracing and microservices observability
Platform & Automation
· Experience with Linux systems administration
· Familiarity with containers and orchestration (Docker, Kubernetes)
· Scripting skills (Bash, Python, or similar)
· Experience integrating observability with CI/CD pipelines
Preferred / Good-to-Have Skills
· Experience with cloud platforms (AWS, Azure, or GCP)
· Infrastructure as Code (Terraform, CloudFormation, etc.)
· Experience working in large-scale distributed systems
· Exposure to service mesh observability (Istio, Linkerd)
· Knowledge of security and compliance monitoring
Soft Skills
· Strong troubleshooting and analytical skills
· Excellent communication skills for cross-team collaboration
· Proactive, ownership-driven mindset