We are looking for a Mid-Level Site Reliability Engineer to improve the reliability, scalability, security, and operational efficiency of our cloud-based systems. You will work closely with software engineers to automate infrastructure and delivery processes, strengthen observability, respond to production incidents, and reduce repetitive operational work.
Key Responsibilities
- Maintain and improve the reliability, availability, performance, and security of production systems.
- Provision and manage cloud infrastructure using infrastructure-as-code tools such as Terraform, Bicep, or CloudFormation.
- Operate and improve containerised workloads running on Kubernetes, managed container platforms, or comparable orchestration services.
- Build and maintain CI/CD pipelines that support safe, repeatable, and automated application and infrastructure deployments.
- Implement and maintain monitoring, structured logging, metrics, distributed tracing, dashboards, and actionable alerts.
- Work with engineering teams to define service-level indicators and service-level objectives for critical services.
- Participate in incident response, investigate production issues, and contribute to blameless post-incident reviews.
- Translate incident findings into concrete improvements to infrastructure, automation, monitoring, documentation, and application resilience.
- Automate repetitive operational activities using appropriate scripting or programming languages.
- Improve system resilience through appropriate scaling, redundancy, health checks, timeouts, retries, backup, and recovery mechanisms.
- Manage application configuration, secrets, service identities, permissions, and environment-specific settings securely.
- Support capacity planning, performance analysis, and cloud-resource optimisation.
- Maintain operational documentation, troubleshooting guides, architectural diagrams, and runbooks.
- Identify operational toil and work with engineering teams to eliminate it through automation or system improvements.
- Assist development teams with production-readiness reviews and help ensure services are observable, deployable, and supportable.
- Participate in a scheduled on-call rotation if required for the assigned product.
- Use AI-assisted tools responsibly to support troubleshooting, automation, documentation, and analysis while validating all generated output.
Must-Have Skills
- 3+ years of professional experience in Site Reliability Engineering, DevOps, cloud infrastructure, systems engineering, or a comparable production-focused role.
- Hands-on experience with at least one major cloud provider and its compute, networking, identity, storage, database, and monitoring services.
- Practical experience with infrastructure-as-code using Terraform, Bicep, CloudFormation, or an equivalent technology.
- Experience working with Docker and containerised applications.
- Practical experience with Kubernetes, a managed container platform, or a comparable application-orchestration environment.
- Experience building, maintaining, or troubleshooting CI/CD pipelines using Azure DevOps, GitHub Actions, GitLab CI, or an equivalent platform.
- Experience with monitoring and observability tools, including logs, metrics, dashboards, alerts, and distributed traces.
- Working understanding of service-level indicators, service-level objectives, availability, latency, error rates, and system saturation.
- Experience investigating production incidents and performing structured root-cause analysis.
- Ability to automate operational work using Python, Bash, PowerShell, Go, or another suitable scripting or programming language.
- Good understanding of Linux systems, processes, permissions, resource usage, and common troubleshooting tools.
- Understanding of networking fundamentals, including DNS, HTTP/HTTPS, TLS, load balancing, firewalls, routing, and private networking.
- Understanding of cloud-security fundamentals, including identity and access management, least privilege, secrets management, and secure configuration.
- Familiarity with relational databases, message queues, caching systems, and the operational concerns associated with them.
- Understanding of backup, recovery, redundancy, scaling, and disaster-recovery principles.
- Ability to troubleshoot issues across infrastructure, applications, networks, containers, databases, and deployment pipelines.
- Ability to communicate clearly during incidents and collaborate effectively with software engineers and other stakeholders.
- Good written and spoken English and the ability to work effectively in a distributed team.
Nice-to-Have Skills
- Hands-on experience with AWS services such as EKS, ECS, EC2, Lambda, RDS, S3, IAM, CloudWatch, Route 53, SQS, SNS, or Secrets Manager.
- Hands-on experience with Microsoft Azure services such as AKS, Container Apps, App Service, Functions, Service Bus, Key Vault, Application Insights, or Azure Monitor.
- Experience administering production Kubernetes clusters and troubleshooting workloads, networking, storage, and resource constraints.
- Familiarity with Helm, Kustomize, Argo CD, Flux, or other Kubernetes and GitOps tooling.
- Experience with observability technologies such as OpenTelemetry, Prometheus, Grafana, Loki, Elasticsearch, Datadog, or Application Insights.
- Experience defining and using SLOs, error budgets, and reliability indicators to support engineering decisions.
- Experience managing on-premises or hybrid-cloud infrastructure.
- Familiarity with incident-management tools and processes, including escalation procedures, post-incident reviews, and follow-up tracking.
- Experience with configuration-management or automation tools such as Ansible.
- Understanding of progressive-delivery strategies such as rolling, blue-green, and canary deployments.
- Experience with vulnerability scanning, dependency security, container-image security, or cloud security-posture management.
- Experience with cloud cost monitoring, capacity planning, and resource optimisation.
- Experience supporting distributed systems, event-driven architectures, or high-availability services.
- Familiarity with .NET/C# applications and the operational characteristics of the .NET runtime.
- Relevant AWS, Azure, Kubernetes, infrastructure, or security certifications..
- Experience working in consultancy, agency, or multi-client delivery environments.