The Investigo Group

The Investigo Group

Hiring Regions We’re excited that you’re interested in joining our team! At the moment, we’re only able to hire applicants who are based in the UK (including Ireland) and the Netherlands. We hope to expand to more locations in the future, so thank you ...

Professional Services
Founded 2023

Description

  • Operate, harden, and extend production OpenShift, OKD, and Kubernetes clusters across on-premises and hybrid environments.
  • Support the migration from VMware to KVM and help modernize the underlying compute and storage layer.
  • Own and improve CI/CD processes across the full lifecycle of platform and application components.
  • Develop and mature GitOps deployment practices using tools such as Argo CD or Flux.
  • Maintain core platform services including identity, ingress, observability, certificate management, service mesh, and container registry capabilities.
  • Build and operate observability across logs, metrics, traces, alerting, SLOs, and error budgets.
  • Improve platform hardening for secure and regulated environments, including network policy, SELinux, image provenance, secret management, and audit controls.
  • Automate repeatable operational tasks using infrastructure and scripting tools such as Ansible, Terraform, Helm, Kustomize, Go, or Python.
  • Lead incident response, support blameless post-mortems, and drive systemic fixes.
  • Partner with networking and security teams on platform integration, segmentation, load balancing, and accreditation evidence.
  • Create and maintain documentation, runbooks, design notes, and operational guidance.
  • Mentor other engineers and act as a senior technical authority across cloud and Kubernetes operations.

Requirements

  • Strong experience running production Kubernetes environments.
  • Strong Linux fundamentals, including systemd, networking, storage, and performance troubleshooting.
  • Experience with at least one Kubernetes distribution such as OKD, OpenShift, vanilla Kubernetes, Rancher, EKS, AKS, or GKE.
  • Solid infrastructure as code experience, including Ansible plus Terraform or equivalent, alongside tools such as Helm and Kustomize.
  • GitOps and CI/CD experience managing full application and component lifecycles, using tools such as Argo CD, Flux, GitHub Actions, or similar.
  • Experience with observability tooling such as Prometheus, Grafana, Elastic Stack/LGTM, or OpenTelemetry.
  • Experience working with identity and access technologies such as OIDC, SAML, SCIM, or Keycloak.
  • Experience with virtualization or infrastructure platforms such as KVM, libvirt, or VMware.
  • Scripting or tooling experience using Go, Python, shell scripting, or similar.
  • Experience working in secure, regulated, or enterprise-scale environments.
  • Strong troubleshooting, problem-solving, and analytical skills.
  • Strong communication skills with the ability to produce clear documentation, runbooks, post-mortems, and technical guidance.
  • Eligible to hold UK SC clearance, including the right to work in the UK and meeting the stated residency requirements.
  • Specific OpenShift or OKD experience, including operators, MachineConfig, or SCCs, is desirable.
  • Service mesh experience such as Istio or Linkerd is desirable.
  • Policy engine experience such as OPA, Gatekeeper, or Kyverno is desirable.
  • Cloud-native application deployment experience using Helm, Terraform, Kustomize, or similar is desirable.
  • Storage experience such as Ceph, Longhorn, or OpenShift Data Foundation is desirable.
  • Networking experience including BGP, VXLAN, Palo Alto, or Juniper technologies is desirable.
  • Software supply chain security experience, including SBOMs, image signing, admission control, or tools such as Sigstore, is desirable.
  • Experience operating AI, ML, or GPU-enabled platforms is desirable.
  • CKA, CKAD, CKS, Red Hat certifications, or equivalent are desirable.
  • Active or recent UK SC clearance is desirable.
  • Recognised open-source contributions to the Kubernetes ecosystem are desirable.
  • Calm, structured, and methodical under pressure.
  • Collaborative working style across platform, development, QA, security, networking, and architecture teams.
  • Strong sense of ownership and accountability.
  • Automation-first mindset with a focus on removing repeatable manual work.
  • Able to influence technical practice through evidence, example, and credibility.
  • Pragmatic and solutions-focused approach to problem solving.
  • Curious about why systems fail, not just how to bring them back online.
  • Comfortable mentoring others and raising the technical capability of those around you.
  • Able to balance reliability, delivery pace, security, and compliance in a regulated environment.

Benefits

  • Private medical health cash plan.
  • 4x life assurance.
  • Generous holiday allowance.
  • Access to continuous learning and development opportunities.
  • Bonus potential based on performance and business-related factors.
  • Discounts on a wide range of products and services.
  • Pension scheme contributions.
  • EV car scheme.
  • Regular pay reviews.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

AI infrastructure Engineer (SRE) Bangalore

Together 1-10 IT Services

Together AI is hiring an AI Infrastructure Engineer (SRE) to keep its user-facing services and production systems reliable, scalable, and available as the company builds next-generation AI infrastructure.

Ansible Kubernetes Machine Learning PagerDuty Terraform
7 hours, 35 minutes ago

Senior Site Reliability Engineer

Omilia 251-1K IT Services

Omilia is hiring a Senior Site Reliability Engineer to operate and improve cloud-based production platforms, observability, and reliability practices across development and engineering teams.

Agile Ansible AWS Bash CentOS Go Grafana Kubernetes MySQL PostgreSQL Prometheus Python Redis TCP/IP Terraform
8 hours, 5 minutes ago

Site Reliability Engineer II

MRSOOL 1K-5K Air Freight & Logistics

Mrsool is hiring an experienced Site Reliability Engineer to help ensure the stability and reliability of its delivery platform while supporting feature delivery and infrastructure growth.

Ansible AWS Azure Chef Docker GCP Go Grafana Java Kubernetes Nagios Prometheus Puppet Python Ruby Terraform
8 hours, 35 minutes ago

Senior Site Reliability Engineer

Latitude AI 501-1000 information technology & services

Latitude AI is hiring a Site Reliability Engineer to help operate and improve the mission-critical systems behind Ford’s autonomous driving platform.

AWS CloudFormation Elasticsearch GCP Go Jaeger Kubernetes Linux Prometheus Python TCP/IP Terraform
1 day, 7 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers