Experience: 8+ Years
Location: Gurugram (On-site)
Employment Type: Full-time
Role Overview
We are looking for a Lead DevOps Engineer with 8+ years of experience to lead the design, implementation, and operations of highly available, scalable, and secure infrastructure across primarily on-premises environments with exposure to AWS Cloud. The ideal candidate will have deep expertise in Rancher Kubernetes Engine (RKE), Kubernetes, Docker, Linux, observability, production operations, and incident management.
As a technical leader, you will drive platform reliability, automation, observability, and operational excellence while mentoring engineers and collaborating closely with development teams to ensure resilient and efficient production systems.
Key Responsibilities
Lead the design, implementation, and management of DevOps and SRE practices across on-premises and AWS environments.
Own the reliability, availability, scalability, and performance of production platforms.
Deploy, manage, and troubleshoot Kubernetes clusters, with strong expertise in RKE (Rancher Kubernetes Engine).
Manage containerized workloads using Docker and Kubernetes.
Build, maintain, and optimize CI/CD pipelines to enable reliable application deployments.
Design and implement observability solutions using OpenTelemetry (OTel), Loki, and other monitoring tools.
Administer and troubleshoot Apache Kafka clusters to ensure reliable event streaming.
Configure and manage Kong API Gateway for API traffic management and security.
Perform Linux system administration, capacity planning, patching, and performance tuning.
Lead incident response, root cause analysis (RCA), and post-incident reviews to improve platform stability.
Automate infrastructure provisioning and operational workflows using Infrastructure as Code (IaC).
Collaborate with development teams to improve application reliability, deployment strategies, and operational readiness.
Mentor junior engineers and establish DevOps/SRE best practices.
Ensure platform security, backup, disaster recovery, and compliance standards are met.
Required Skills
8+ years of experience in DevOps, SRE, or Platform Engineering.
Strong expertise in Linux administration and troubleshooting.
Hands-on experience with Kubernetes, particularly RKE (Rancher Kubernetes Engine).
Strong experience with Docker and containerized environments.
Experience managing on-premises infrastructure (mandatory) with exposure to AWS Cloud.
Strong operational knowledge of Apache Kafka.
Hands-on experience with OpenTelemetry (OTel) and Loki for observability.
Experience configuring and managing Kong API Gateway.
Strong troubleshooting skills across infrastructure, Kubernetes, networking, and distributed systems.
Experience with CI/CD tools such as Jenkins, GitLab CI, GitHub Actions, or similar.
Hands-on experience with Infrastructure as Code (Terraform or Ansible).
Good understanding of networking concepts, load balancers, DNS, SSL/TLS, and security best practices.
Experience in production support, incident management, and SRE methodologies.
Preferred Skills
Experience with Prometheus and Grafana.
Knowledge of GitOps tools such as Argo CD.
Experience working in hybrid cloud environments.
Scripting experience using Bash or Python.
Kubernetes certification (CKA/CKAD) or AWS certification is a plus.