Job Description – DevOps / Infrastructure Engineer
Experience: 3+ Years
Location: Gurgaon/ On-site
Employment Type: Full-time
Role Overview
We are looking for a DevOps / Infrastructure Engineer with 3+ years of hands-on experience in AWS, Kubernetes, Terraform, Ansible, CI/CD, monitoring, and infrastructure management. The candidate will be responsible for maintaining and automating infrastructure, supporting production environments, troubleshooting issues, and ensuring system reliability and security.
Key Responsibilities
Manage and maintain AWS infrastructure, Linux servers, and production environments.
Handle server provisioning, configuration, patching, upgrades, and troubleshooting.
Build and maintain CI/CD pipelines for application deployments.
Manage containerized applications using Docker and Kubernetes.
Support Kubernetes deployments, troubleshooting, scaling, and upgrades.
Automate infrastructure provisioning using Terraform.
Automate server configuration and operational tasks using Ansible.
Implement monitoring, logging, and alerting using Prometheus, Grafana, Datadog, ELK, and Loki.
Monitor infrastructure and application health and troubleshoot production issues.
Develop Python scripts for automation and operational activities.
Follow Infosec and infrastructure security best practices, including access control, vulnerability remediation, and server hardening.
Participate in production incident management and root-cause analysis.
Support infrastructure and application upgrades with minimal downtime.
Collaborate with development and security teams to improve deployment, automation, and operational processes.
Required Skills
3+ years of experience in DevOps / Infrastructure / Cloud Engineering
Strong hands-on experience with AWS
Kubernetes and Docker
Terraform and Ansible
CI/CD pipelines
Linux / Server Management
Python scripting
Prometheus and Grafana
ELK / Loki
Datadog
Infrastructure monitoring and troubleshooting
Basic to good understanding of Infosec and cloud security
Experience with production environments, upgrades, and incident management