Staff Reliability Engineer (Full Stack)

1 month, 3 weeks ago
Full-time
Lead
DevOps and Infrastructure
Feeld

Feeld

Feeld is a modern dating app that caters to open-minded individuals seeking fulfilling relationships. It provides a space for curious and open-minded humans to explore intimacy, embrace desires, and connect with like-minded people. Through Feeld, users...

Family Services
51-250
Founded 2014

Description

  • Own reliability outcomes for critical backend services and their integration with React Native mobile clients.
  • Lead incident response by coordinating mitigation, diagnosing root causes, communicating status, and driving resolution.
  • Build and improve monitoring and observability through dashboards, alerts, tracing, and logging.
  • Run blameless post-incident reviews and turn learnings into durable fixes, runbooks, automation, and process updates.
  • Improve engineering safety through guardrails, safer migrations, feature-flag practices, rollout strategies, and resilience patterns.
  • Partner with product, design, QA, and engineering to align delivery plans with operational risk and reliability needs.
  • Strengthen documentation and onboarding materials such as architecture notes, service ownership docs, runbooks, and working guides.
  • Mentor engineers through pairing, code reviews, incident shadowing, and coaching on production ownership.
  • Collaborate across squads to improve production ownership, reliability, and backend-to-mobile integration patterns.

Requirements

  • Significant experience building and operating production backend systems at scale, including debugging distributed systems and performance issues.
  • Strong TypeScript/Node.js backend experience, or equivalent, with comfort working across services and APIs.
  • Proven incident response leadership experience, including on-call participation, triage, mitigation, and root-cause analysis with follow-through.
  • Solid observability skills with practical experience in logging, metrics, tracing, dashboards, and actionable alerting.
  • Experience collaborating with mobile teams on backend-to-mobile integration concerns such as API compatibility, releases, and feature flags.
  • Demonstrated staff-level IC leadership through design reviews, technical direction, documentation, and cross-team alignment.
  • React Native experience or a strong understanding of mobile architecture patterns and release constraints is preferred.
  • AWS or similar cloud experience, plus familiarity with infrastructure as code, CI/CD, and production tooling is preferred.
  • Experience designing reliability programs such as SLOs, error budgets, incident processes, and operational excellence improvements is preferred.
  • Experience with PostgreSQL, Redis, and performance tuning in high-traffic systems is preferred.
  • Experience in a high-growth environment where prioritization and pragmatic trade-offs are essential is preferred.

Benefits

  • Flexible working hours.
  • Unlimited paid time off.
  • Fully remote work.
  • Home office budget.
  • Learning and development budget.
  • On-demand therapy sessions and mental health support via Spill.
  • In-person meetups.
  • Transparent, equitable compensation with a baseline freedom salary of £60,000 GBP per year for roles below that amount.
  • Estimated market-competitive total cash compensation of £100,000 to £130,000 GBP, depending on geographic location.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Lead Site Reliability Engineer (Performance & Scalability) | Contract | Remote US

Tech Holding 51-250 Internet Software & Services

Tech Holding is hiring a contract Lead Site Reliability Engineer to assess and improve the performance, reliability, and scalability of its platform for current and future demand.

CI/CD
15 hours, 30 minutes ago

Site Reliability Engineer

Coforge 10K-50K IT Services

Coforge is hiring a remote Site Reliability Engineer to help build and operate reliable cloud platforms and production systems for teams across Costa Rica, Peru, Colombia, and Bolivia.

AWS Azure C# CI/CD Datadog GCP Go Grafana Java Kubernetes Linux Microservices New Relic OpenTelemetry Prometheus Python Splunk Terraform TypeScript
1 day, 15 hours ago

AMS:US:SL:Service Reliability Engineer:Lead

Thoughtworks 10K-50K Professional Services

Thoughtworks is hiring a Service Reliability Engineer to improve the reliability, resilience, and performance of client infrastructure and production systems.

Ansible CircleCI CloudFormation ELK Stack GitLab GitOps Go Grafana Jaeger Java Jenkins Kubernetes Nomad Prometheus Python Ruby Shell Scripting Terraform Zipkin
1 day, 16 hours ago

Sr. Staff Platform/Data Reliability Engineer, Databricks (R5537)

Bitly 51-250 Internet Software & Services

Shield AI is seeking a Sr. Staff Platform / Data Reliability Engineer to make its Databricks platform reliable, secure, scalable, and operationally mature for enterprise and regulated use.

CI/CD Databricks Git
3 days, 16 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers