Arbor

Arbor

Arbor is the leading cloud MIS provider in the UK, empowering schools and MATs to collaborate effectively, save time, and enhance pupil achievement through centralized data management and insightful analytics.

IT Services
51-250

Description

  • Define and guide system architecture, balancing speed, scalability, maintainability, and security.
  • Champion reliability and performance by ensuring systems are observable and meet agreed SLOs.
  • Lead root cause analysis and help improve incident response processes and frameworks.
  • Drive automation initiatives to reduce operational toil and improve system efficiency.
  • Uphold coding standards, promote automated testing, and ensure production readiness standards.
  • Lead technical estimation, feasibility assessments, release planning, and post-release reviews.
  • Mentor and coach engineers through feedback, knowledge sharing, and technical guidance.
  • Collaborate with Product Managers, Engineering Managers, and engineers to align technical direction with product strategy.
  • Communicate complex technical concepts clearly to technical and non-technical stakeholders.

Requirements

  • Extensive professional experience in SRE, DevOps, or Platform Engineering on complex, scalable systems.
  • Extensive expertise with AWS and distributed cloud architectures.
  • Proven experience operating platforms serving a high volume of requests, around 1000 requests per second.
  • Advanced proficiency with Terraform and configuration management tools.
  • Strong programming skills in Python, Go, or a similar language for automation and tooling.
  • Deep experience with monitoring and observability platforms such as DataDog or Prometheus, plus incident/problem management.
  • Expert understanding of distributed systems, microservices, and resilience patterns.
  • Hands-on experience with containerization and orchestration technologies such as Docker, Kubernetes, or ECS.
  • Practical experience building and maintaining CI/CD pipelines for automated deployments.
  • Demonstrated ability to mentor and support the growth of fellow engineers.
  • Experience with chaos engineering and reliability testing is a bonus.
  • Knowledge of security best practices and compliance frameworks is a bonus.
  • Background in agile and lean methodologies such as Scrum or Kanban is a bonus.
  • Contributions to open-source projects or the SRE community are a bonus.
  • Visa sponsorship is not available for this role.

Benefits

  • Salary of £80,000 - £90,000.
  • Remote working.
  • 32 days holiday including Bank Holidays, made up of 25 days annual leave plus 7 company-wide days.
  • Life assurance at 3x annual salary.
  • Private dental insurance with Bupa.
  • Salary sacrifice pension provided by Scottish Widows.
  • Enhanced maternity and adoption leave of 20 weeks full pay, and paternity leave of 6 weeks full pay.
  • Access to wellbeing support, including mindfulness, mental health first aid training, Calm, Bippit, and AIG Smart Health with 24/7 virtual GP, counselling, and health checks.
  • Flexible working arrangements discussed to suit individual needs.
  • Dedicated professional development budget for CPD courses, upskilling resources, and professional memberships.
  • Volunteer with a charity of your choice for one day each year.
  • Dog-friendly offices.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

OMS Application Engineer

Warner Music Group is seeking an OMS Application Engineer to maintain, upgrade, and support the technical systems powering its global commerce ecosystem, with a focus on reliability, scalability, and long-term stability.

AWS CI/CD GitHub Actions Java Microservices
3 days, 7 hours ago

[Job-32080] Senior SRE / Cloud Engineer (Pessoa Engenheira de Plataforma), Brazil

CI&T 5K-10K Internet Software & Services

A CI&T busca uma pessoa Engenheira de Plataforma Sênior para projetar e operar a infraestrutura OCI que sustenta agentes de IA, garantindo alta disponibilidade e baixa latência.

Generative AI Kafka Kubernetes Terraform
3 days, 7 hours ago

Sr. Sustaining and Forward Deployed Engineer

Abacus Insights 51-250 Insurance

Abacus Insights is seeking a Senior Site Reliability Engineer – Forward Deployed to operate and improve its AWS- and Databricks-based healthcare data platform while resolving complex production issues and supporting customers.

Apache Spark AWS CI/CD Databricks Kubernetes Python Snowflake
4 days, 8 hours ago

Site Reliability Engineer

Raydar Professional Services

A Site Reliability Engineer at a remote conversational AI company will build and operate the CI/CD, developer platforms, observability, and agent infrastructure supporting reliable voice-agent products at enterprise scale.

CI/CD Helm Kubernetes Node.js Python Terraform TypeScript
5 days, 8 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers