Role Overview
We are looking for a Senior SRE & Platform Engineer to design, build, and maintain our deployment automation systems. In this role, you will bridge the gap between software development and systems engineering. You will be responsible for creating seamless CI/CD pipelines that empower developers to ship code rapidly, while simultaneously ensuring our production environments are secure, highly available, and scalable.
Core Responsibilities
- Kubernetes Platform Ownership: Architect, provision, and maintain secure, multi-region Kubernetes clusters (e.g., EKS, GKE, or AKS) using Infrastructure as Code (IaC).
- Deployment Automation & GitOps: Design, implement, and optimize scalable CI/CD pipelines and GitOps workflows (e.g., Argo CD, Flux) to automate software delivery in containers.
- Observability & Performance: Establish robust monitoring, logging, and tracing frameworks (e.g., Prometheus, Grafana, OpenTelemetry) to track cluster health and system latency.
- Reliability Engineering: Define, measure, and enforce Service Level Objectives (SLOs) and Error Budgets. Run regular chaos engineering drills and disaster recovery testing.
- Toil Elimination: Write clean, maintainable automation scripts and operators (Go, Python, Bash) to replace repetitive operational tasks and auto-remediate known infrastructure issues.
Required Technical Skills
- Containerization & Orchestration: Advanced, production-level experience managing Kubernetes clusters, custom controllers, Helm charts, and container runtimes (Docker/containerd).
- Infrastructure as Code: Proven expertise with Terraform, OpenTofu, or Pulumi for cloud infrastructure provisioning.
- CI/CD & GitOps: Strong experience with deployment tooling such as Argo CD, Flux, GitHub Actions, GitLab CI, or Jenkins.
- Cloud Platforms: Hands-on experience with at least one major cloud provider (AWS, GCP, or Azure), focusing on networking, IAM, and managed services.
- Programming & Scripting: Proficiency in Python or Go for building infrastructure tooling, alongside robust Bash scripting skills.
Soft Skills & Qualifications
- 5–7+ years of experience in an SRE, DevOps, or Platform Engineering role managing cloud-native production environments.
- Strong systems mindset with a deep understanding of Linux internals, networking (DNS, TCP/IP, BGP), and security best practices.
- Exceptional communication skills with the ability to collaborate with product development teams to improve application architecture for reliability.