About the Role
We are seeking an experienced GCP DevOps Specialist to design, build, and manage cloud infrastructure on Google Cloud Platform (GCP). This role is core to our platform engineering function — you'll own the reliability, scalability, and security of our GCP-based environments while enabling teams building AI/ML products to ship faster and safer. Strong, hands-on Google Cloud Platform expertise is central to this role, along with Kubernetes and Linux systems skills.
Key Responsibilities
Architect, deploy, and manage cloud infrastructure on Google Cloud Platform (GCP) — Compute Engine, GKE, VPC, Cloud IAM, Cloud Storage, Load Balancing, and Cloud Monitoring
Own end-to-end GCP infrastructure provisioning using Infrastructure-as-Code (Terraform preferred)
Design and manage Kubernetes clusters (GKE) for scalable, highly available application delivery
Build and maintain CI/CD pipelines integrated with GCP-native tooling (Cloud Build, Artifact Registry) and third-party tools
Perform Linux system troubleshooting — performance tuning, networking, and OS-level issue resolution across GCP compute environments
Support infrastructure for AI/ML workloads, including GPU provisioning and ML pipeline deployment on GCP
Implement monitoring, logging, and alerting for GCP workloads using Cloud Monitoring, Prometheus, Grafana
Drive GCP cost optimization, auto-scaling, and resource efficiency initiatives
Implement GCP security best practices — IAM policies, network security, compliance, and audit readiness
Collaborate cross-functionally with engineering and data science teams to support production systems running on GCP
Required Skills & Experience
4–6 years of hands-on DevOps / Platform Engineering experience
Strong, demonstrable expertise in Google Cloud Platform (GCP) — this is the core requirement for the role, including:
Compute Engine, GKE (Google Kubernetes Engine)
VPC networking, Cloud Load Balancing
IAM & Cloud security controls
Cloud Monitoring / Cloud Logging
Kubernetes experience is mandatory — deployment, scaling, and troubleshooting
Strong Linux troubleshooting skills (system, network, and performance issues)
Exposure to AI/ML infrastructure or ML workflow support
CI/CD pipeline experience (Cloud Build, Jenkins, GitHub Actions, or GitLab CI)
Infrastructure-as-Code experience, preferably Terraform on GCP
Scripting skills in Python or Bash for automation
Solid understanding of Docker and container orchestration best practices
Good to Have
Google Cloud Professional certification (e.g., Professional Cloud DevOps Engineer / Professional Cloud Architect)
Experience with MLOps tools and GPU-based workloads on GCP
Familiarity with GCP-native monitoring/observability tooling