We are looking for a highly skilled and certified Senior DevOps Engineer to join our growing infrastructure team. In this role, you will be responsible for designing, building, and maintaining robust CI/CD pipelines, containerised workloads, cloud infrastructure, and observability platforms that underpin our business-critical financial applications.
The ideal candidate brings deep AWS expertise, strong Linux and scripting skills, hands-on experience with Kubernetes orchestration (EKS/ECS), Infrastructure-as-Code, and a forward-looking mindset toward GenAI-augmented DevOps practices.
Architect, provision, and manage production-grade AWS environments across multiple accounts and regions.
Lead EKS (Elastic Kubernetes Service) cluster management: node group scaling, upgrades, RBAC, pod security policies, and network policies.
Design and manage ECS (Elastic Container Service) workloads using Fargate and EC2 launch types.
Configure and manage AWS services: VPC, EC2, RDS, S3, IAM, Route53, ALB/NLB, ACM, CloudWatch, SNS, SQS, Lambda, Secrets Manager, Parameter Store, WAF, Shield, and API Gateway.
Implement and maintain multi-account AWS organisation strategies with SSO, SCPs, and guardrails.
Ensure high availability, disaster recovery (DR), and business continuity for all production workloads.
Design and maintain Docker images and multi-stage Dockerfiles optimised for security and performance.
Define Helm charts for Kubernetes application deployments across dev, staging, and production environments.
Implement Kubernetes best practices: resource limits, health probes, HPA/VPA, pod disruption budgets, and rolling deployments.
Manage container image repositories and lifecycle policies via ECR.
Build, maintain, and improve CI/CD pipelines using Jenkins (declarative & scripted pipelines).
Integrate automated testing, code quality gates (SonarQube), vulnerability scanning, and artifact management into pipelines.
Manage Git repositories, branching strategies (GitFlow / trunk-based), and code review workflows.
Implement blue-green, canary, and rolling deployment strategies to support zero-downtime releases.
Collaborate with development teams on shift-left security and DevSecOps practices.
Own end-to-end Terraform codebases: module design, state management, remote backends (S3 + DynamoDB), and workspace strategies.
Implement Terraform CI/CD with plan/apply automation, drift detection, and policy-as-code (OPA/Sentinel or Checkov).
Maintain version-controlled IaC with documented runbooks and change management processes.
Design and manage a unified observability stack using Prometheus, Grafana, and New Relic.
Build Grafana dashboards for infrastructure KPIs, application performance, and business metrics.
Define alerting rules, SLOs, SLIs, and error budgets aligned with business-critical service SLAs.
Implement distributed tracing and log aggregation (ELK / CloudWatch Logs / Loki).
Conduct regular capacity planning reviews and performance tuning exercises.
Manage AWS API Gateway (REST & HTTP APIs): throttling, WAF integration, authorisers, usage plans, and versioning.
Design and maintain secure VPC architecture: subnets, NACLs, security groups, VPC peering, Transit Gateway, and PrivateLink.
Implement mTLS, API security policies, and rate-limiting for all external-facing services.
Deep expertise in Linux (RHEL/CentOS/Amazon Linux/Ubuntu) administration and hardening.
Perform kernel tuning, filesystem management, process management, and performance profiling.
Implement OS-level security hardening aligned with CIS benchmarks.
Manage SSH, PAM, sudoers, and user access controls in production environments.
Write production-quality Bash scripts for system automation, alerting, log rotation, patching, and health-check routines.
Develop Python scripts for AWS automation (boto3), data processing, operational tooling, and custom Prometheus exporters.
Build internal CLI tools and runbook automations to reduce MTTR and eliminate toil.
Actively explore and integrate GenAI tools into DevOps workflows — AI-assisted code reviews, incident summarisation, runbook generation, and intelligent alerting.
Evaluate and prototype LLM-powered internal tools (Copilot for DevOps, ChatOps bots, AIOps capabilities).
Stay current on AI/ML infrastructure trends and support data/ML teams with MLOps tooling where applicable.
Experiment with prompt engineering for automation use cases and share findings for team knowledge sharing.
Implement secrets management using AWS Secrets Manager, HashiCorp Vault, or equivalent.
Conduct regular vulnerability assessments, patching cycles, and security audits of cloud and container environments.
Ensure compliance with NBFC regulatory requirements (data residency, audit trails, access logging).
Integrate SAST/DAST tools and container image scanning (Trivy/Snyk) into CI/CD pipelines.
Act as the escalation point for production incidents impacting business-critical applications.
Lead post-incident reviews (PIRs), root cause analyses (RCAs), and corrective action implementation.
Define and maintain runbooks, SOPs, and disaster recovery playbooks.
Drive SRE practices including chaos engineering and game days for resilience validation.
Mentor junior and mid-level DevOps engineers; conduct code reviews and enforce quality standards.
Partner with development, QA, security, and product teams to align DevOps practices with organisational goals.
Present infrastructure roadmaps and cost optimisation proposals to technology leadership.
Contribute to vendor evaluations, RFPs, and toolchain decisions.
AWS (Expert Level) — AWS Certification required:
AWS Certified Solutions Architect – Associate or Professional
AWS Certified DevOps Engineer – Professional (preferred)
AWS Certified SysOps Administrator (added advantage)
Kubernetes & Containers: EKS, ECS, Docker, Helm (3+ years hands-on)
CI/CD: Jenkins (advanced pipeline development), Git (GitFlow, PR workflows)
Infrastructure as Code: Terraform (advanced — modules, workspaces, remote state)
Monitoring & Observability: Prometheus, Grafana, New Relic
API Management: AWS API Gateway (REST & HTTP)
Linux: Advanced — RHEL/Amazon Linux/Ubuntu, bash scripting, system hardening
Scripting: Bash (advanced), Python (intermediate–advanced with boto3)
Networking: VPC design, DNS, load balancing, TLS/SSL, firewall rules
Security: IAM, KMS, Secrets Manager, RBAC, CIS hardening, vulnerability scanning
Hands-on exposure or experimentation with GenAI tools (GitHub Copilot, ChatGPT API, LangChain)
Understanding of LLMOps or MLOps concepts is a strong plus
Ability to leverage AI for code generation, documentation, and operational automation
Strong analytical and problem-solving mindset under high-pressure production scenarios
Excellent written and verbal communication — ability to translate complex technical concepts for non-technical stakeholders
Ownership mentality with a proactive approach to identifying and resolving risks
Team player with experience working in Agile/Scrum delivery models
Minimum 4 years in a dedicated DevOps / SRE / Cloud Infrastructure role
Minimum 3 years working on AWS production environments at scale
Minimum 2 years with Kubernetes (EKS preferred) in production
Experience in BFSI / NBFC / FinTech domain is a strong advantage
Experience managing business-critical, 24x7 production applications