Primary Responsibilities:
-
Trend Analysis & Problem Identification
-
Identify recurring incident patterns, anomalies, and signs of alert
fatigue that may indicate deeper systemic issues
-
Collaborate with L2/L3 teams to review telemetry data and recommend
improvements to alert thresholds, rules, and policies
-
Provide insights that support proactive issue prevention, noise
reduction, and overall monitoring refinement
-
Platform Management & Optimization
-
Develop, update, and maintain dashboards that reflect real‑time system
health, performance metrics, and service behavior
-
Hands-on experience in setting up end-to-end application monitoring,
including configuration of key metrics, thresholds, Service Levels,
and Synthetic monitors using monitoring platforms.
-
Support the ongoing adoption and optimization of Dynatrace
(Zabbix/ monitoring tool)
, enhancing dashboarding and visualization capabilities for cloud and
on‑prem observability
-
Assist in routine platform checks, ensuring monitoring tools remain
accurate, stable, and aligned with business and operational
requirements
-
Collaboration & Coordination
-
Work closely with application teams, SRE groups, and infrastructure
operations during incident triage, investigations, and routine
monitoring reviews
-
Ensure clear, timely, and effective communication with stakeholders
during service-impacting events, providing status updates and context
as needed
-
Organizing the work for the team, including planning, task breakdown,
and ensuring clarity of priorities
-
Provides structured, timely updates to leadership on progress, risks,
blockers, team capacity, and delivery timelines
-
Ensures adherence to engineering best practices, drives operational
excellence, and maintains accountability for team delivery outcomes
-
Operational Excellence
-
Support platform stability and availability through adherence to
lifecycle maintenance, patching schedules, and vulnerability
management processes
-
Contribute to the improvement of monitoring workflows, alert routing
logic, runbook effectiveness, and incident management practices
-
Innovation & AI Enablement
-
Assist in exploring and adopting AI-driven capabilities that improve
observability, automate root‑cause identification, and reduce manual
effort
-
Contribute to internal knowledge sharing by documenting best
practices, playbooks, AI reference materials, and usage guidelines
(e.g., Copilot tips)
-
Collaboration & Leadership Support
-
Partner with cross-functional teams to align monitoring practices with
evolving business needs and operational priorities
-
Drive end to end delivery of monitoring initiatives—requirements
gathering, planning, execution oversight, and delivery validation
-
Coordinate cross‑team dependencies, ensure timelines are met, and
proactively remove blockers for the team
-
Provide subject‑matter support for ITSM processes including incident,
problem, and change management discussions
-
Comply with the terms and conditions of the employment contract, company
policies and procedures, and any and all directives (such as, but not
limited to, transfer and/or re-assignment to different work locations,
change in teams and/or work shifts, policies in regards to flexibility of
work benefits and/or work environment, alternative work arrangements, and
other decisions that may arise due to the changing business environment).
The Company may adopt, vary or rescind these policies and directives in
its absolute discretion and without any limitation (implied or otherwise)
on its ability to do so
Qualifications - External
Required Qualifications:
-
Bachelor’s in computer science, Information Technology, or related field
or equivalent experience
-
6+ years in Site Reliability Engineering or Observability/Monitoring
engineering roles
-
5+ years hands-on with monitoring/observability tools: Zabbix
-
4+ years of scripting experience any scripting
-
4+ year scripting with Python and Bash or PowerShell for automation
-
2+ years with Azure (architecture fundamentals, observability in
cloud-native and lift‑and‑shift contexts)