Job Summary
We are looking for an experienced Senior Site Reliability Engineer (SRE) with 7–8 years of experience in designing, implementing, and maintaining highly available, scalable, and reliable cloud infrastructure on Google Cloud Platform (GCP).
The ideal candidate should have strong hands-on experience with GCP, Kubernetes, Linux, Infrastructure as Code, CI/CD, observability, automation, and production operations. The role will involve improving system reliability, performance, scalability, and operational efficiency while working closely with development and platform teams.
Key Responsibilities
* Design, deploy, and manage highly available and scalable infrastructure on GCP.
* Own production infrastructure, reliability, availability, performance, and operational excellence.
* Build and maintain GKE/Kubernetes clusters and containerized workloads.
* Implement and manage CI/CD pipelines for application and infrastructure deployments.
* Automate repetitive operational tasks using Python, Bash, or similar scripting languages.
* Develop and maintain Infrastructure as Code using Terraform.
* Implement monitoring, logging, alerting, and observability using Google Cloud Monitoring, Cloud Logging, and other tools.
* Define and monitor SLIs, SLOs, and SLAs for critical services.
* Participate in production incident management, troubleshooting, root-cause analysis, and post-incident reviews.
* Perform capacity planning, performance tuning, and scalability assessments.
* Implement high-availability, disaster recovery, backup, and business-continuity strategies.
* Identify reliability risks and proactively implement solutions to reduce system downtime.
* Work with development teams to improve application reliability and deployment processes.
* Establish and improve operational runbooks, automation, and self-healing mechanisms.
* Support security and compliance requirements across cloud infrastructure.
* Participate in on-call/production support activities when required.
Required Technical Skills
GCP
* Strong hands-on experience with Google Cloud Platform.
* Experience with:
* GKE
* Compute Engine
* VPC
* IAM
* Cloud Load Balancing
* Cloud Storage
* Cloud SQL
* Pub/Sub
* Cloud Monitoring & Logging
* Secret Manager
* Good understanding of GCP networking, IAM, security, and cost optimization.
Kubernetes & Containers
* Strong hands-on experience with Kubernetes and GKE.
* Docker/containerization.
* Kubernetes networking, deployments, services, ingress, configmaps, secrets, RBAC, and troubleshooting.
* Experience with Helm and Kubernetes operators is preferred.
Infrastructure as Code
* Strong experience with Terraform.
* Experience designing reusable Terraform modules and managing infrastructure through Git-based workflows.
CI/CD & DevOps
* Strong understanding of CI/CD principles.
* Hands-on experience with tools such as:
* GitLab CI/CD
* Jenkins
* GitHub Actions
* ArgoCD
* Experience implementing automated deployment and rollback strategies.
Observability & Reliability
* Experience with monitoring, logging, tracing, and alerting.
* Hands-on experience with Prometheus, Grafana, OpenTelemetry or equivalent tools.
* Strong understanding of:
* SLIs
* SLOs
* SLAs
* Error Budgets
* Incident Management
* RCA
Linux & Networking
* Strong Linux administration and troubleshooting skills.
* Good understanding of:
* TCP/IP
* DNS
* HTTP/HTTPS
* Load Balancing
* SSL/TLS
* Network troubleshooting
* Firewall and security concepts.
Scripting & Automation
* Strong scripting experience in Python and/or Bash.
* Ability to build automation tools and operational scripts.
Good to Have
* GCP Professional Cloud DevOps Engineer certification.
* Experience with GitOps and ArgoCD.
* Experience with service mesh such as Istio.
* Experience with SRE frameworks and practices.
* Experience with chaos engineering and resilience testing.
* Experience with distributed systems and microservices architecture.
* Experience with FinOps/cloud cost optimization.
* Exposure to security best practices and DevSecOps.
* Experience working in 24x7 production environments.
Required Soft Skills
* Strong analytical and problem-solving skills.
* Excellent troubleshooting and debugging ability.
* Good communication and documentation skills.
* Ability to work independently and take ownership of production systems.
* Strong incident management and decision-making skills.
* Ability to collaborate effectively with Development, Security, Product, and Platform teams.
Education
Bachelor’s degree in Computer Science, Engineering, Information Technology, or a related field is preferred.