DevOps Engineer
We are seeking an experienced DevOps Engineer with 3-7 years of hands-on
experience in cloud infrastructure, infrastructure as code, and observability
solutions. You will be responsible for designing, deploying, and maintaining
robust cloud infrastructure while implementing comprehensive monitoring,
logging, and orchestration solutions.
Position Overview
We are seeking an experienced DevOps Engineer with 3-7 years of hands-on
experience in cloud infrastructure, infrastructure as code, and observability
solutions. You will be responsible for designing, deploying, and maintaining
robust cloud infrastructure while implementing comprehensive monitoring,
logging, and orchestration solutions.
Key Responsibilities
-
Design, deploy, and manage cloud infrastructure using AWS, Azure, or GCP
with a focus on scalability, reliability, and cost optimization.
-
Develop and maintain Infrastructure as Code (IaC) using Terraform across
multi-cloud environments.
-
Implement and manage comprehensive observability solutions using LangFuse,
Prometheus, and Grafana for real-time system monitoring.
-
Configure and manage CI/CD pipelines using GitHub Actions for automated
testing, building, and deployment.
-
Utilize Harness for advanced deployment orchestration, feature flags, and
pipeline management.
-
Develop automation scripts in Python or Bash Shell to streamline operational
tasks and reduce manual overhead.
-
Establish and maintain monitoring alerts, dashboards, and runbooks for
proactive incident management.
-
Troubleshoot and resolve production issues while maintaining high system
availability.
-
Collaborate with software development teams to ensure smooth deployment
processes and operational best practices.
Required Qualifications
-
3-7 years of professional DevOps or Infrastructure Engineering experience.
-
Hands-on experience with at least one cloud platform: AWS, Azure, or Google
Cloud Platform (GCP).
-
Strong proficiency with Terraform for Infrastructure as Code, including
module development and state management.
-
Demonstrated experience setting up and managing LangFuse for LLM
observability and monitoring.
-
Hands-on experience with Prometheus for metrics collection and time-series
data.
-
Experience with Grafana for creating dashboards, alerts, and visualization
of monitoring data.
-
Practical experience with Harness for deployment orchestration and pipeline
automation.
-
Proficiency with GitHub Actions for building and deploying CI/CD pipelines.
-
Strong scripting skills in Python or Bash Shell for automation and
infrastructure management.
-
Understanding of containerization technologies (Docker, Kubernetes) is
highly valued.
-
Solid understanding of networking, security best practices, and compliance
requirements.