Cloud Infrastructure
-
Design, deploy, and manage scalable, highly available, and secure AWS
infrastructure.
-
Architect cloud environments capable of supporting enterprise-scale
applications and AI workloads.
-
Optimize infrastructure for performance, reliability, scalability, and cost.
-
Implement high availability, disaster recovery, backup, and failover
strategies.
-
Design multi-environment infrastructure (Development, QA, UAT, Production).
Kubernetes & Container Platform
-
Design, deploy, and manage production-grade Kubernetes clusters using
Amazon EKS
.
-
Optimize Kubernetes workloads for high availability and resource
utilization.
-
Configure namespaces, RBAC, network policies, autoscaling, ingress
controllers, and service meshes where applicable.
-
Troubleshoot Kubernetes networking, scheduling, storage, and performance
issues.
-
Manage rolling deployments, blue-green deployments, and canary releases.
AWS Services
Strong hands-on experience with:
-
Amazon EKS
-
Amazon EC2
-
Auto Scaling Groups
-
Elastic Load Balancer (ALB/NLB)
-
Amazon S3
-
Amazon RDS
-
AWS Lambda
-
Amazon ECR
-
Amazon CloudWatch
-
IAM
-
Route 53
-
VPC
-
NAT Gateway
-
Security Groups
-
AWS WAF
-
AWS Secrets Manager
-
Systems Manager (SSM)
-
CloudFront
-
EventBridge
-
SNS
-
SQS
CI/CD & DevOps Automation
-
Design and implement end-to-end CI/CD pipelines.
-
Automate application deployments across multiple environments.
-
Implement infrastructure automation and GitOps practices.
-
Build deployment strategies with minimal downtime.
-
Integrate automated testing, security scanning, and quality gates into CI/CD
pipelines.
Experience with:
-
GitHub Actions
-
Jenkins
-
GitLab CI
-
ArgoCD
Infrastructure as Code
Develop and manage infrastructure using:
-
Terraform
-
AWS CloudFormation
-
Kubernetes YAML
AI & LLM Infrastructure
-
Deploy and manage Small Language Models (SLMs) and Large Language Models
(LLMs) in production environments.
-
Build scalable inference infrastructure for AI workloads.
-
Configure GPU-enabled Kubernetes nodes for model serving.
-
Optimize CPU and GPU utilization for AI inference.
-
Manage model deployments, scaling, versioning, and monitoring.
-
Support vector databases and AI inference services.
-
Work closely with AI/ML engineers to optimize model performance and
infrastructure costs.
Database Infrastructure & Performance
-
Deploy and manage Amazon RDS databases.
-
Monitor and optimize database performance.
-
Implement backup, recovery, and replication strategies.
-
Tune database configurations for high-throughput applications.
-
Monitor slow queries, indexing strategies, and connection pooling.
-
Collaborate with engineering teams on database performance optimization.
Monitoring & Observability
Implement monitoring and observability using:
-
CloudWatch
-
Prometheus
-
Grafana
-
ELK / OpenSearch
-
Loki
Responsibilities include:
-
Infrastructure monitoring
-
Application monitoring
-
Log aggregation
-
Alerting
-
Capacity planning
-
Incident response
Security & Compliance
-
Implement AWS security best practices.
-
Design secure IAM policies and access controls.
-
Manage secrets and encryption.
-
Perform infrastructure hardening.
-
Ensure compliance with:
-
Participate in security audits and vulnerability remediation.
-
Maintain audit logs and infrastructure documentation.
Cost Optimization
-
Continuously optimize AWS infrastructure costs.
-
Right-size EC2 instances and EKS node groups.
-
Optimize storage and networking costs.
-
Implement Savings Plans and Reserved Instances where appropriate.
-
Optimize GPU utilization for AI workloads.
-
Monitor cloud spending and recommend cost-saving initiatives.
Required Qualifications
-
Bachelor's or Master's degree in Computer Science, Information Technology,
or a related field.
-
5–8 years of hands-on experience
in DevOps, Cloud Engineering, or Platform Engineering.
-
Strong experience designing and managing production AWS environments.
-
Extensive experience with Kubernetes and Amazon EKS.
-
Experience managing enterprise-scale cloud infrastructure.
-
Proven experience automating deployments and infrastructure management.
Required Technical Skills
Cloud Platforms
-
Amazon Web Services (AWS)
AWS Services
-
Amazon EC2
-
Amazon EKS
-
Amazon ECS
-
Amazon RDS
-
Amazon S3
-
Lambda
-
ECR
-
CloudFront
-
IAM
-
Route 53
-
VPC
-
CloudWatch
-
Systems Manager
-
WAF
-
Secrets Manager
-
SNS
-
SQS
-
EventBridge
Containers & Orchestration
-
Docker
-
Kubernetes
-
Amazon EKS
-
Helm
-
Kubernetes Networking
-
Ingress Controllers
-
Horizontal & Vertical Pod Autoscaling
Infrastructure as Code
-
Terraform
-
CloudFormation
-
Helm
-
Kustomize
CI/CD
-
GitHub Actions
-
Jenkins
-
GitLab CI
-
ArgoCD
Databases
-
Amazon RDS
-
PostgreSQL
-
MySQL
-
Redis
Experience with:
-
Performance tuning
-
Replication
-
Backup & recovery
-
Connection pooling
-
Query optimization
AI Infrastructure
Experience deploying and managing:
-
LLMs and SLMs
-
GPU-based inference workloads
-
NVIDIA GPU infrastructure
-
CUDA-enabled environments (preferred)
-
Hugging Face models
-
vLLM, Ollama, or similar inference frameworks
-
Model serving and autoscaling
Monitoring & Logging
-
Prometheus
-
Grafana
-
CloudWatch
-
ELK/OpenSearch
-
Loki
Security & Compliance
Strong understanding of:
-
SOC 2
-
HITRUST
-
HIPAA
-
IAM
-
RBAC
-
Network Security
-
Encryption
-
Secrets Management
-
Vulnerability Management
Preferred Qualifications
-
AWS Certified Solutions Architect – Professional or Associate.
-
AWS Certified DevOps Engineer – Professional.
-
Certified Kubernetes Administrator (CKA) or Certified Kubernetes Application
Developer (CKAD).
-
Experience with AI platforms, MLOps, or GPU infrastructure.
-
Experience deploying high-availability, multi-tenant SaaS applications.
-
Familiarity with service mesh technologies (Istio or Linkerd) is a plus.