Loading open roles
Loading open roles
Loading role

GirnarSoft · posted 1 month ago
| Position / Designation | Site Reliability Engineer (SRE) – AI & Cloud Infrastructure |
| Qualification | Bachelor’s degree in a relevant field |
| Years of Experience | 5–8 Years |
| Permanent / Contract (If contract, period ?) | c ontract - 6 month |
| Office / Remote / Hybrid | r emote |
| Number of post | 1 |
| Gender | MALE / FEMALE |
| Industry Background | IT |
| Annual CTC / Salary | As per company norms |
| Selection Process | 2 (Technical & HR Rounds) |
| Role & Responsibilities |
· Design, implement, and manage highly available, scalable, and secure cloud infrastructure on AWS. · Build and maintain an end-to-end observability platform using Open Telemetry, Grafana, Datadog, CloudWatch, and related tools. · Implement AIOps capabilities, including: · LLM-assisted incident triage · AI-powered root cause analysis · ML-driven forecasting and anomaly detection · Intelligent alert correlation and noise reduction · Lead production incident management, on-call response, postmortems, and Root Cause Analysis (RCA). · Automate operational workflows using Infrastructure as Code (Terraform/CloudFormation) and CI/CD pipelines. · Drive infrastructure rightsizing, capacity planning, utilization analysis, and cloud cost optimization. · Build dashboards, SLOs, SLIs, and error budgets to improve service reliability. · Develop automation scripts using Python, Bash, or Go to eliminate manual operational tasks. · Monitor application and infrastructure health while ensuring high uptime and service performance. · Collaborate with Development, DevOps, Security, Platform Engineering, and Product teams across multiple time zones. · Establish operational best practices for monitoring, incident response, disaster recovery, and resilience engineering. · Maintain Linux-based production systems and troubleshoot OS, networking, storage, and performance issues. |
| Skills & Qualification |
· 5–8 years of experience in Site Reliability Engineering, DevOps, or Cloud Infrastructure Engineering. · Strong experience with AWS services including EC2, ECS/EKS, Lambda, VPC, IAM, CloudWatch, RDS, Route 53, S3, and Auto Scaling. · Hands-on experience with Open Telemetry, Grafana, Datadog, Prometheus, or similar monitoring platforms. · Strong knowledge of Linux administration, networking, system performance tuning, and troubleshooting. · Experience with Infrastructure as Code using Terraform or CloudFormation. · Proficiency in scripting using Python, Bash, or Go. · Experience with Kubernetes and containerized workloads. · Strong understanding of CI/CD pipelines (GitHub Actions, Jenkins, GitLab CI, etc.). · Experience leading incident management, production support, and RCA processes. · Knowledge of SRE principles including SLIs, SLOs, and Error Budgets. · Experience implementing monitoring, logging, alerting, and observability frameworks. · Strong analytical, troubleshooting, and communication skills. Preferred Qualifications · Experience building or implementing AIOps solutions. · Exposure to Large Language Models (LLMs) for operational automation. · Experience with machine learning-based forecasting or anomaly detection. · Hands-on experience administering Adobe Experience Manager (AEM). · Experience managing Cloudflare CDN, WAF, DNS, and caching strategies. · Knowledge of FinOps, cloud cost optimization, and capacity planning. · AWS Solutions Architect, DevOps Engineer, or Kubernetes certifications are a plus.
|
| Joining Date | Immediate joiners |