Flip App

Flip App

Flip is the employee app reshaping workplace communication by empowering every employee with a digital workspace for effective communication and workflow management.

Internet Software & Services
51-250
Founded 2018

Description

  • Own critical reliability domains end-to-end within the Platform Squad.
  • Drive technical direction and architectural decisions for the platform.
  • Help evolve cloud infrastructure on Azure and Kubernetes for high throughput and high availability.
  • Define and improve the platform’s resilience strategy, including scaling, zero-downtime deployments, rollback mechanisms, and disaster recovery.
  • Improve the observability stack built around Loki, Grafana, Tempo, and Mimir.
  • Reduce infrastructure toil by making the IaC platform more self-service for engineering teams.
  • Lead platform-related major incidents and drive blameless post-mortems.
  • Coach teammates, run RFCs and design reviews, and mentor engineers within the squad.
  • Partner with the squad to shape the platform roadmap and direction.

Requirements

  • 5+ years of hands-on experience as an SRE, Platform Engineer, DevOps Engineer, Infrastructure Engineer, Cloud Engineer, or Backend Engineer with a strong infrastructure focus.
  • Proven track record of building and operating high-throughput, highly available production systems.
  • Deep production-level experience with Kubernetes on any hyperscaler.
  • Strong experience with modern observability stacks such as Prometheus, Mimir, VictoriaMetrics, Dash0, Loki, or ELK, plus a clear point of view on SLIs, SLOs, and error budgets.
  • Solid software development skills in Go, strongly preferred because the IaC runs on Pulumi in Go, or Python.
  • Hands-on experience with Infrastructure as Code tools such as Pulumi, OpenTofu, or Terraform, plus GitOps tools such as ArgoCD and CI/CD pipeline design.
  • Demonstrated ability to lead complex infrastructure initiatives from design to production, including writing RFCs and driving architecture decisions.
  • Experience mentoring engineers and raising the technical bar within a team.
  • Comfortable owning major incidents end-to-end and turning learnings into systemic change.
  • Strong communication skills and business-fluent English.
  • Willingness to participate in on-call rotations.
  • Preferred: experience rolling out production-ready API gateways with Gateway API such as Envoy Gateway.
  • Preferred: experience operating multi-cluster service meshes such as Cilium, Linkerd, or Istio.
  • Preferred: experience deploying and maintaining Kubernetes Operators such as Strimzi or CNPG.
  • Preferred: experience operating highly available PostgreSQL in production.

Benefits

  • Remote-first work with flexibility to work from home.
  • Occasional team events, workshops, or meetings in the Berlin or Stuttgart offices with plenty of notice.
  • E-Gym-Wellpass membership covered by the company.
  • Job bike leasing.
  • Regular team events and culture days.
  • Option to work abroad within the European Union.
  • Relaxed working atmosphere with highly motivated and committed colleagues.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Site Reliability Engineer

PENN Entertainment 10K-50K Hotels, Restaurants & Leisure

PENN Entertainment is seeking a Senior Site Reliability Engineer to build and operate the cloud infrastructure supporting large-scale sports betting and media platforms across regulated jurisdictions.

Argo CD AWS Bash CI/CD Datadog GCP GitHub Actions GitOps Go Helm Kubernetes Linux PostgreSQL Python Shell Scripting Terraform
17 hours, 14 minutes ago

Staff Site Reliability Engineer, Ads

Reddit 1K-5K Internet Software & Services

Reddit is hiring a Staff Site Reliability Engineer to provide technical leadership for reliability, scalability, and operational excellence across its advertising infrastructure and revenue-critical systems.

Apache Spark ClickHouse GCP Go Kafka Kubernetes
17 hours, 44 minutes ago

Sr Lead Network Reliability Engineer

Coupa Software 1K-5K Internet Software & Services

Coupa is hiring a Sr. Lead Network Development Engineer to scale and operate its global SaaS platform’s cloud networking infrastructure through automation, reliability engineering, and technical leadership.

Ansible AWS Azure Chef DNS Fortinet Go Java Kubernetes Linux Python Ruby TCP/IP Terraform TLS
1 day, 17 hours ago

Senior Monitoring/Observability Architect

Makpar 51-250 Internet Software & Services

Makpar is seeking a Senior Monitoring/Observability Architect to lead enterprise monitoring strategy, architecture, and implementation guidance for a large federal government program.

Datadog Splunk
1 day, 17 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers