Must have
Self-managed production experience — not RDS/Atlas/MSK — with at least two of MySQL, PostgreSQL, MongoDB
Linux and TCP/IP troubleshooting at the command line: sockets, routing, DNS, TLS, packet capture
Kubernetes day-2 operations: real diagnosis of Pending, CrashLoop, probe, PVC and DNS failures; Helm values discipline
Has built or upgraded a cluster, not only consumed one — any distribution counts, including a vendor-managed on-prem platform
Ansible roles or playbooks that someone other than the author runs
Has been on-call for something that mattered, and got to a root cause
Deploy and upgrade Java/Spring Boot services on Kubernetes clusters inside customer datacentres, using Helm with per-customer value sets.
Manage the Kubernetes clusters themselves — installation and version upgrades, node lifecycle and capacity, control-plane and etcd health, certificate rotation, CNI, ingress and storage classes — both on clusters we build and on enterprise distributions the customer provides.
Provision, tune and operate self-managed databases on Linux VMs — MySQL, PostgreSQL and MongoDB — including replication, backup, verified point-in-time restore, and version upgrades.
Provision and operate the middleware layer on VMs — Redis (Sentinel), Apache Kafka and RabbitMQ — including clustering, authentication and capacity planning.
Run OpenSearch as the centralised application log store: shard and heap sizing, index lifecycle and retention policies, cluster health.
Automate the full infrastructure build with Ansible — reusable roles plus per-customer inventories, so the same automation installs the stack at every customer without being forked.
Own day-2 Kubernetes operations: rollouts, probe and resource tuning, persistent storage, ingress and TLS, and diagnosis of pod, scheduling and networking failures.
Troubleshoot production issues across the VM, network, Kubernetes and JVM boundaries — determine whether a fault is infrastructure or application, and drive root-cause analysis.
Debug connectivity problems at the command line: routing, firewalls, DNS, MTU, TLS chains and certificate trust, packet capture.
Manage certificates, truststores and private container registry (Harbor) content for restricted and air-gapped environments.
Monitor the platform using Prometheus, Grafana and OpenSearch — define what gets alerted on, not only what gets dashboarded.
Write Python and Shell automation for provisioning, health checks and routine operational tasks.
Apply security and access hygiene appropriate to regulated customer environments — least-privilege access, secret handling, network segmentation, audit readiness.
Produce and maintain runbooks and handover documentation so a customer's own operations team can run what we install.
Work directly with customer infrastructure, network and security teams during deployment and issue resolution, on-site where required.
Collaborate with development and QA teams to plan and execute releases.
Support periodic out-of-hours deployment and upgrade windows at customer sites.