Staff Reliability Engineer (Full Stack)

3 weeks, 4 days ago
Full-time
Lead
DevOps and Infrastructure
Feeld

Feeld

Feeld is a modern dating app that caters to open-minded individuals seeking fulfilling relationships. It provides a space for curious and open-minded humans to explore intimacy, embrace desires, and connect with like-minded people. Through Feeld, users...

Family Services
51-250
Founded 2014

Description

  • Own reliability outcomes for critical backend services and their integration with React Native mobile clients.
  • Lead incident response by coordinating mitigation, diagnosing root causes, communicating status, and driving resolution.
  • Build and improve monitoring and observability through dashboards, alerts, tracing, and logging.
  • Run blameless post-incident reviews and turn learnings into durable fixes, runbooks, automation, and process updates.
  • Improve engineering safety through guardrails, safer migrations, feature-flag practices, rollout strategies, and resilience patterns.
  • Partner with product, design, QA, and engineering to align delivery plans with operational risk and reliability needs.
  • Strengthen documentation and onboarding materials such as architecture notes, service ownership docs, runbooks, and working guides.
  • Mentor engineers through pairing, code reviews, incident shadowing, and coaching on production ownership.
  • Collaborate across squads to improve production ownership, reliability, and backend-to-mobile integration patterns.

Requirements

  • Significant experience building and operating production backend systems at scale, including debugging distributed systems and performance issues.
  • Strong TypeScript/Node.js backend experience, or equivalent, with comfort working across services and APIs.
  • Proven incident response leadership experience, including on-call participation, triage, mitigation, and root-cause analysis with follow-through.
  • Solid observability skills with practical experience in logging, metrics, tracing, dashboards, and actionable alerting.
  • Experience collaborating with mobile teams on backend-to-mobile integration concerns such as API compatibility, releases, and feature flags.
  • Demonstrated staff-level IC leadership through design reviews, technical direction, documentation, and cross-team alignment.
  • React Native experience or a strong understanding of mobile architecture patterns and release constraints is preferred.
  • AWS or similar cloud experience, plus familiarity with infrastructure as code, CI/CD, and production tooling is preferred.
  • Experience designing reliability programs such as SLOs, error budgets, incident processes, and operational excellence improvements is preferred.
  • Experience with PostgreSQL, Redis, and performance tuning in high-traffic systems is preferred.
  • Experience in a high-growth environment where prioritization and pragmatic trade-offs are essential is preferred.

Benefits

  • Flexible working hours.
  • Unlimited paid time off.
  • Fully remote work.
  • Home office budget.
  • Learning and development budget.
  • On-demand therapy sessions and mental health support via Spill.
  • In-person meetups.
  • Transparent, equitable compensation with a baseline freedom salary of £60,000 GBP per year for roles below that amount.
  • Estimated market-competitive total cash compensation of £100,000 to £130,000 GBP, depending on geographic location.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Site Reliability Engineer

Counterpart Health 51-200 hospital & health care

Counterpart Health is hiring a Senior Site Reliability and Infrastructure Engineer to support and evolve the technology platform behind its primary care tool and maintain reliable infrastructure for domestic and international workloads.

AWS Azure CI/CD Containerd DNS Docker GCP Go gRPC Helm Kubernetes Linux Load Balancing Prometheus Python Shell Scripting TCP/IP
14 hours, 59 minutes ago

Senior Test Platform & Reliability Engineer - Star Trek Fleet Command

Scopely 1K-5K Internet Software & Services

Scopely is hiring a Senior Test Platform & Reliability Engineer in Ireland to build validation, reliability, and developer enablement platforms for Star Trek Fleet Command’s large-scale live-service backend systems.

AWS Bash CI/CD Docker GitLab Go Python Terraform
15 hours, 14 minutes ago

Senior Software Engineer - Databases, SRE | Canada | Remote

Grafana 1K-5K IT Services

Grafana Labs is hiring a Senior Software Engineer for its remote SRE team to improve reliability and operability of Grafana Cloud database services for high-SLA customers across AWS, GCP, and Azure.

AWS Azure GCP Go Helm Java Kubernetes Linux Microservices Python Terraform
1 day, 14 hours ago

Senior Site Reliability Engineer

Semios 51-250 Food Products

Semios Group is hiring a Senior Site Reliability Engineer to help scale, secure, and improve the reliability of its global agricultural technology platform.

AWS Azure Bash Buildkite CI/CD Datadog Docker Envoy GCP Git GitHub GitHub Actions GitLab Go Jenkins Kubernetes Linux NATS New Relic Prometheus Python Ruby Splunk Terraform
1 day, 15 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers