Loading open roles
Loading open roles
Loading role

Production Support
Missions
We are seeking a highly skilled DevOps / Site Reliability Engineer (SRE) with strong expertise in monitoring, logging, tracing, and observability platforms . The ideal candidate will design, implement, and operate scalable observability and monitoring solutions to ensure high availability, performance, and reliability of business-critical systems.
This role requires hands-on experience with modern observability stacks including Prometheus, Grafana, ELK, OpenTelemetry , and related tools, along with strong operational and automation skills.
Experience
5-8 years in DevOps / SRE / Observability roles
Key Responsibilities
DevOps & SRE Responsibilities
· Design, implement, and maintain highly available, scalable, and reliable systems
· Apply SRE principles such as SLIs, SLOs, error budgets, and incident management
· Automate infrastructure provisioning, monitoring, and alerting processes
· Participate in on-call rotations and lead incident triage and root cause analysis (RCA)
· Improve system reliability through proactive monitoring and performance tuning
Monitoring & Observability
· Build and manage end-to-end observability platforms covering:
o Metrics
o Logs
o Traces
· Design and implement monitoring frameworks for cloud-native and distributed systems
· Define alerting strategies to reduce noise and improve mean time to detect (MTTD) and recover (MTTR)
· Create and maintain dashboards, alerts, and runbooks for operational teams
Tooling & Platform Management
· Implement and manage observability tools including:
o Metrics : Prometheus, Thanos
o Visualization : Grafana
o Logging : Elasticsearch, Logstash, Kibana (ELK), Fluentd
o Tracing : Jaeger, Tempo
o Telemetry : OpenTelemetry
· Integrate observability tools with CI/CD, cloud platforms, and Kubernetes environments
· Optimize storage, retention, and performance of observability data
Tooling & Platform Management
· Implement and manage observability tools including:
o Metrics : Prometheus, Thanos
o Visualization : Grafana
o Logging : Elasticsearch, Logstash, Kibana (ELK), Fluentd
o Tracing : Jaeger, Tempo
o Telemetry : OpenTelemetry
· Integrate observability tools with CI/CD, cloud platforms, and Kubernetes environments