Key Responsibilities
AWS Cloud Infrastructure
Provision, configure, harden and maintain Amazon EC2 instances in accordance with approved architecture and security standards.
Design and operate snapshot-based backup schedules, retention policies and recovery procedures; verify backup completion and conduct periodic restoration tests.
Configure and administer VPC components, subnets, route tables, internet and NAT gateways, security groups and network access controls.
Deploy and manage load balancers, health checks, target groups and related availability and traffic-routing configurations.
Create, review and tune AWS WAF rules to protect applications from common web threats while minimising false positives.
Administer Amazon RDS, Route 53, EventBridge, CloudWatch and Lambda services for database hosting, DNS, event automation, monitoring and operational remediation.
Monitor resource health, availability, capacity, performance and cost signals; recommend corrective or preventive improvements.
Maintain accurate cloud asset inventories, architecture records, access reviews and operating documentation.
Networking, DNS and Perimeter Security
Manage DNS records and domain mappings, validate name resolution, and coordinate changes with application owners and external providers.
Install, renew and troubleshoot SSL/TLS certificates, certificate chains and secure protocol configurations.
Implement and review IP allowlisting and access restrictions with appropriate approvals, traceability and periodic validation.
Configure and support Sophos security appliances, Wi-Fi routers, firewall policies, VPN or network access settings, and related firmware updates.
Configure and monitor static WAN IP connectivity, track link status and coordinate timely resolution with internet service providers.
Review firewall and router logs, investigate anomalies, and maintain secure baseline configurations and backup copies of device settings.
Linux and Application Platform Administration
Provide expert-level administration of Ubuntu and other Linux environments, including users, groups, permissions, processes, services, storage, packages, logs and system performance.
Plan and execute operating-system patching, security updates and controlled reboots for deployed Ubuntu servers, including pre-checks, rollback planning and post-change validation.
Install, configure, secure, monitor and troubleshoot Nginx, Apache Tomcat, SFTP and SSH services.
Apply secure configurations for authentication, ciphers, ports, file permissions, service accounts, logging and remote administration.
Troubleshoot server, network and application issues using logs, system utilities and structured root-cause analysis.
Automate repeatable operational tasks where practical and maintain reusable runbooks and scripts under version control.
Monitoring, Alerting and Incident Response
Install, configure and administer Nagios monitoring servers, agents, checks, host and service definitions, dashboards and notification rules.
Manage monitoring access, contact groups, escalation paths, maintenance windows and alert routing.
Monitor infrastructure, applications, databases, network links and critical services; tune thresholds to produce actionable alerts.
Respond to incidents, coordinate technical recovery, communicate status and document timelines, root causes and corrective actions.
Review availability and capacity trends and proactively address recurring faults, saturation risks and single points of failure.
Backup, Data Recovery and Business Continuity
Own periodic backup validation, data-recovery exercises and business continuity planning for critical services.
Define recovery procedures, dependencies, responsibilities, communication paths and evidence requirements in collaboration with business and application stakeholders.
Manage deployment and testing of recovery processes against approved recovery time and recovery point objectives.
Record test outcomes, gaps and corrective actions; drive closure of findings and keep recovery documentation current.
Support disaster-recovery events and ensure that restoration activities protect data integrity, security and business priorities.
Security, VAPT Remediation and Compliance
Analyse infrastructure-related VAPT findings, assess operational impact and remediate vulnerabilities within agreed timelines.
Harden server and application response headers, including removal of unnecessary information and implementation of appropriate security headers.
Review and securely configure CORS policies in coordination with application teams, allowing only required origins, methods and headers.
Apply secure configuration baselines for Linux, Nginx, Tomcat, SSH, SFTP and MySQL; maintain evidence of remediation and validation.
Support access reviews, audit requests, vulnerability scans and technical risk assessments.
Ensure changes follow documented approval, testing, rollback and evidence-retention processes.
MySQL Database Administration
Administer MySQL users, roles and privileges in line with least-privilege and segregation-of-duties principles.
Plan and execute MySQL upgrades, security patches and maintenance with backups, compatibility checks, rollback plans and post-change verification.
Install, configure and monitor MySQL primary-replica replication, including replication status, lag, errors, capacity and recovery procedures.
Harden database configurations, network exposure, authentication, logging, encryption and backup settings.
Support database availability, performance troubleshooting, restoration exercises and incident resolution in partnership with application teams.
Azure and Cross-Cloud Support
Provide basic administration and operational support for Microsoft Azure resources, identities, networking, monitoring and access controls.
Apply consistent security, documentation, monitoring and change-management practices across AWS, Azure and on-premises environments.
Assist with cross-cloud troubleshooting, service comparisons and workload transition planning when required.
Technical Leadership and Governance
Lead day-to-day infrastructure operations, prioritise work and provide hands-on guidance to engineers and support personnel.
Translate business requirements into secure, supportable technical solutions and clearly communicate risks, dependencies and trade-offs.
Review technical changes, enforce operational standards and ensure appropriate testing, approvals, documentation and rollback readiness.
Coordinate with software engineering, security, vendors and business stakeholders during releases, incidents, audits and continuity exercises.
Maintain standard operating procedures, architecture diagrams, asset records, configuration baselines, credentials-management processes and knowledge articles.
Promote automation, preventive maintenance, continuous improvement and effective knowledge transfer.
Required Experience and Qualifications
Minimum one year of hands-on experience performing the infrastructure, cloud, Linux, networking, monitoring, database, security and recovery activities described in this job description.
Practical experience administering production or business-critical AWS and Linux environments.
Demonstrated experience with EC2, snapshots, VPC, security groups, load balancers, WAF, RDS, Route 53, EventBridge, CloudWatch and Lambda.
Strong working knowledge of TCP/IP, DNS, HTTPS, SSL/TLS, routing, firewalls, VPN concepts, IP allowlisting and network troubleshooting.
Expert-level Linux administration skills, with strong command-line proficiency and a sound understanding of system security and performance.
Hands-on experience with Nagios, Nginx, Tomcat, SFTP, SSH and MySQL administration.
Experience remediating VAPT findings, including response-header, CORS, operating-system, database and service-configuration issues.
Experience planning and testing backup, recovery and business continuity procedures.
Basic operational knowledge of Microsoft Azure.
Ability to prepare clear runbooks, incident reports, change records, architecture notes and recovery documentation.
Core Competencies
Ownership: Takes responsibility for service health, security, follow-through and documentation.
Technical judgement: Diagnoses complex issues, evaluates risk and selects practical, supportable solutions.
Leadership: Sets priorities, reviews work, mentors colleagues and coordinates action during incidents.
Security mindset: Applies least privilege, secure defaults, timely patching and evidence-based remediation.
Communication: Explains technical issues, impact and recovery status clearly to technical and non-technical stakeholders.
Continuous improvement: Identifies recurring problems and improves automation, monitoring, resilience and operating processes.
Preferred Qualifications
Bachelor’s degree or diploma in Computer Science, Information Technology, Engineering or a related discipline, or equivalent practical experience.
Relevant AWS, Linux, networking, security, database, IT service management or Azure certification.
Experience with infrastructure automation, configuration management, scripting, source control or CI/CD practices.
Familiarity with service-management processes such as incident, problem, change, release and capacity management.
Key Performance Indicators
Availability and stability of critical infrastructure and services.
Backup completion, restoration-test success and business continuity exercise outcomes.
Timely patching and closure of vulnerabilities and VAPT findings.
Monitoring coverage, alert quality, incident response time and recurrence reduction.
Successful infrastructure changes with minimal unplanned impact and effective rollback readiness.
Accuracy and currency of runbooks, inventories, access records and architecture documentation.
Security of cloud, network, Linux, application-platform and database configurations.