Experience
10+ years in Cloud Infrastructure, DevOps, SRE, or Platform Engineering, with strong hands-on experience in Google Cloud Platform (GCP) and Google Kubernetes Engine (GKE).
Role Overview
We are looking for an experienced GCP DevOps Architect to design, implement, and operate highly scalable, secure, resilient, and cost-efficient cloud platforms on GCP.
The ideal candidate should have strong expertise in GCP, GKE, Kubernetes, DevOps, CI/CD, Infrastructure as Code, observability, security, and high-availability architecture, with experience building platforms capable of handling 5–10 million users / high-volume production traffic.
The candidate will be responsible for defining cloud architecture, DevOps standards, deployment strategies, reliability engineering practices, and platform automation.
Key Responsibilities
GCP & Cloud Architecture
Design highly available and scalable cloud architectures on GCP.
Architect platforms capable of supporting 5–10 million users and high-volume traffic.
Design multi-region / multi-zone architectures for high availability and disaster recovery.
Select and optimize appropriate GCP services based on scalability, reliability, performance, and cost.
Design networking architecture including VPC, subnets, Cloud Load Balancing, Cloud NAT, Cloud DNS, Private Service Connect, firewall policies, and hybrid connectivity.
Define cloud architecture standards, reference architectures, and engineering best practices.
GKE & Kubernetes
Strong hands-on experience designing and managing GKE production clusters.
Design scalable GKE architectures including:
Regional and private GKE clusters
Node pools and workload isolation
Cluster autoscaling
Horizontal/Vertical Pod Autoscaling
Workload Identity
Ingress and Gateway architecture
Network policies
Pod security
Optimize Kubernetes workloads for performance, availability, and cost.
Define Kubernetes deployment, upgrade, backup, and disaster recovery strategies.
Troubleshoot complex production issues involving Kubernetes, networking, compute, and application performance.
DevOps & CI/CD
Design and implement enterprise-grade CI/CD pipelines.
Strong experience with tools such as GitHub/GitLab, Jenkins, Cloud Build, Argo CD, and/or other CI/CD platforms.
Implement GitOps-based deployment strategies where appropriate.
Automate build, test, security scanning, deployment, rollback, and release processes.
Implement progressive delivery strategies such as:
Blue/Green deployments
Canary releases
Rolling deployments
Establish DevOps standards across development and operations teams.
Infrastructure as Code
Strong hands-on experience with Terraform.
Build reusable Terraform modules and infrastructure frameworks.
Automate provisioning and configuration of GCP infrastructure.
Implement infrastructure versioning, state management, policy controls, and automated validation.
Experience with configuration management and automation tools is desirable.
Scalability & Reliability
Architect systems for millions of users and high concurrent traffic.
Design autoscaling strategies across compute, Kubernetes, databases, and networking layers.
Implement SRE practices including:
SLIs/SLOs/SLAs
Error budgets
Capacity planning
Reliability engineering
Incident management
Performance engineering
Conduct architecture reviews, scalability assessments, and production readiness reviews.
Design fault-tolerant systems with appropriate RTO/RPO targets.
Monitoring & Observability
Design comprehensive monitoring and observability solutions using Google Cloud Operations Suite / Cloud Monitoring / Cloud Logging and tools such as Prometheus, Grafana, OpenTelemetry, and Datadog.
Implement infrastructure, application, Kubernetes, and business-level monitoring.
Establish centralized logging, metrics, tracing, alerting, and dashboards.
Analyze production performance and identify bottlenecks.
Security
Implement GCP security best practices across infrastructure and Kubernetes.
Strong understanding of:
IAM
Service Accounts
Workload Identity
Secret Manager
KMS
VPC Service Controls
Organization Policies
Security Command Center
Container/image security
Implement least-privilege access and secure CI/CD pipelines.
Integrate vulnerability scanning and security controls into the DevOps lifecycle.
Cost Optimization
Monitor and optimize GCP infrastructure costs.
Optimize GKE compute, node pools, autoscaling, storage, networking, and logging costs.
Establish cloud FinOps practices and cost governance.
Identify opportunities for capacity optimization without compromising reliability.
Required Technical Skills
Must Have:
Strong GCP expertise
Strong GKE / Kubernetes expertise
Strong Terraform / IaC
Advanced CI/CD and DevOps
GCP networking and security
Linux and container technologies
Production experience with highly scalable systems
Experience supporting 5–10 million users or equivalent high-volume traffic
High Availability and Disaster Recovery architecture
Monitoring, logging, and observability
Strong troubleshooting and incident-management skills
Good to Have:
Google Cloud Professional Cloud Architect certification
Google Cloud Professional Cloud DevOps Engineer certification
Service mesh experience such as Istio
Argo CD / GitOps
Prometheus / Grafana / OpenTelemetry
Anthos / GKE Enterprise
FinOps experience
Experience with Kafka, Redis, PostgreSQL/MySQL, NoSQL platforms
Experience with API Gateway / Apigee
Experience with microservices architecture
Leadership & Soft Skills
Strong architectural and problem-solving capabilities.
Ability to communicate complex cloud architecture to both technical and non-technical stakeholders.
Experience mentoring DevOps, SRE, and cloud engineering teams.
Ability to lead architecture decisions and establish engineering standards.
Strong ownership of production reliability and operational excellence.
Ability to work effectively with application, security, data, and infrastructure teams.