Loading open roles
Loading open roles
Loading role

GirnarSoft · posted 4 months ago
Position: SRE developer
Location: Bangalore
Mode: Hybrid (3 days/week from office)
Budget: 1.7 LPM + GST
Experience Required: 5 to 9 years
Notice Period : Immediate
Shift Timings: 6 AM to 3 PM
Onboarding Process Notes: Internal technical round > L1 with client > L2
with end client > Selection
Must-Have Skills
Skill Required Depth
Linux based systems: Ownership
Production Troubleshooting & Observability (Incident RCA, Log Analytics,
Cloud Monitoring): Ownership
Python or Bash: Hands-on
Google Kubernetes Engine (GKE) (Deployments, scaling, basic cluster
troubleshooting): Hands-on
Automation & IaC (Terraform state management and modules): Exposure
Definitions
Ownership - should have owned and be responsible for a part of the project
where the core skill required is mentioned one. He should be able to describe
the project and how he delivered on his specific piece. If required he also
mentored a few juniors on the deliverables. He should be able to describe what
code he wrote, what were the production level challenges faced etc. Can answer
scenario questions on this on an expert level with multiple level probing.
Handson - was working under supervision of someone else but was responsible for something where the core skill required is mentioned on. Should be able to describe the project, his specific role and what were his specific responsibilities. He should be able to describe what code he wrote, what were the production level challenges faced etc. Can answer scenario questions on an average level with 1-2 level probing on specific concepts.
Exposure - was working in a team where someone else was responsible for the skill set in question but the person has seen it in action. He knows the issues that happened with respect to the skill and what did the team do. He is not able to solve scenario level questions but knows typically how to collaborate with other team members on solving the problem.
Knowledge - knows conceptually about how things work and can answer a few questions. Cannot answer anything in scenario based questions.
Detailed JD received from client:
CME Group is seeking a Consultant to help, build, operate and scale systems in our Markets portfolio. Markets SREs work on products and applications related to CME's Globex trading platform. Our systems deliver an exceptional combination of low-latency performance and rock-solid reliability to seamlessly handle the world's busiest trading days.
The successful candidate will work alongside senior engineers to learn how we observe, monitor, automate, and improve Production service reliability and act as a mentor to junior colleagues. He/she will have a keen interest in SRE and enjoy the cut-and-thrust of operating Production systems. They will be a strong communicator, and may have previously worked in an SRE role, a software engineering role or a systems engineering role.
Key Responsibilities:
- Collaborate with senior engineers and product teams to ensure requirements
are mutually understood, planned carefully and implemented safely
- Lead discussions for own work and present solution options and proposals
- Participate in incident response and management — engages with urgency in
live incidents, takes ownership for minor incidents, ensures system recovery
and contributes to post-mortems afterwards
- Participate in on-call rotation
- Identify toil and reduce through automation
- Contribute to DR and systems resiliency testing & improvements
- Contribute own ideas and reliability improvement suggestions to the Product
backlog
- Support the migration of markets applications to Google Cloud Platform (GCP)
- Act as a mentor to L2 and L1 SRE colleagues
What We're Looking for:
- Experience with Linux-based systems
- Experience with Cloud-based platform(s) — Google Cloud Platform, GCE, and/or
GKE a bonus
- Understanding of application architectures and messaging protocols
- Competent programming/scripting skills (Python, Bash, etc.)
- Strong problem-solving and analytical abilities
- Excellent communication and teamwork skills
- Eagerness to learn and adapt in a fast-paced trading environment
Desirable:
- Experience with metrics & monitoring, OpenTelemetry, Splunk, Prometheus,
Grafana, etc.
- Experience and knowledge of working with distributed systems
- Experience with Kubernetes
- Knowledge of networking (HTTP/TCP/UDP/IP)
- Experience in Financial markets
- Experience working in an agile environment