Block

Block

Block is a company that consists of Square, Cash App, Spiral, TIDAL, TBD, and foundational teams. They are focused on economic empowerment by creating tools to expand access to the economy. Square helps sellers run and grow businesses, Cash App redefin...

Capital Markets
10K-50K
Founded 2009

Description

  • Build and extend platforms to improve system reliability.
  • Work toward company-wide reliability goals across multiple platforms and organizations.
  • Standardize reliability tools across teams and systems.
  • Triage, coordinate, and lead stabilization efforts for sev 0–1 incidents.
  • Serve as primary on-call for Tier 0 services and maintain structured escalation paths.
  • Lead incident command, mitigation, and escalation during high-severity events.
  • Drive platform-wide reliability improvements, shared operational tooling, and deploy-safety patterns.
  • Use AI-driven systems to improve signal detection, reduce noise, and accelerate root cause analysis.
  • Design and implement safe deployment patterns such as progressive delivery, automated rollback, and guardrails.
  • Improve observability, incident detection and response, and operational workflows through automation.

Requirements

  • 5+ years of software development experience.
  • Experience running production on-call for high-availability systems.
  • Strong incident management skills, including structured triage, mitigation under pressure, and blameless postmortems.
  • Fluency with CI/CD pipelines, progressive rollout strategies, and rollback automation.
  • Monitoring and observability expertise, including tuning alerts for uptime, error rates, latency regression, and resource exhaustion.
  • Familiarity with AI-driven tooling for observability, incident analysis, or automation.
  • A mindset that naturally uses AI to accelerate problem-solving and reduce toil.
  • Demonstrated technical initiative and leadership on previous backend or platform-focused projects.
  • Ability to create and maintain evidence-based maturity assessments using trailing 90-day data windows.
  • Comfort with vendor and dependency management, including maintaining validated escalation contacts reachable within 5 minutes.

Benefits

  • Remote work.
  • Medical insurance.
  • Flexible time off.
  • Retirement savings plans.
  • Modern family planning benefits.
  • A globally distributed work environment with collaboration across multiple time zones.
  • Reasonable accommodations during the recruitment process for disabled applicants.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Site Reliability Engineer

Coforge 10K-50K IT Services

Coforge is hiring a remote Site Reliability Engineer to help build and operate reliable cloud platforms and production systems for teams across Costa Rica, Peru, Colombia, and Bolivia.

AWS Azure C# CI/CD Datadog GCP Go Grafana Java Kubernetes Linux Microservices New Relic OpenTelemetry Prometheus Python Splunk Terraform TypeScript
5 hours, 39 minutes ago

AMS:US:SL:Service Reliability Engineer:Lead

Thoughtworks 10K-50K Professional Services

Thoughtworks is hiring a Service Reliability Engineer to improve the reliability, resilience, and performance of client infrastructure and production systems.

Ansible CircleCI CloudFormation ELK Stack GitLab GitOps Go Grafana Jaeger Java Jenkins Kubernetes Nomad Prometheus Python Ruby Shell Scripting Terraform Zipkin
6 hours, 9 minutes ago

Sr. Staff Platform/Data Reliability Engineer, Databricks (R5537)

Bitly 51-250 Internet Software & Services

Shield AI is seeking a Sr. Staff Platform / Data Reliability Engineer to make its Databricks platform reliable, secure, scalable, and operationally mature for enterprise and regulated use.

CI/CD Databricks Git
2 days, 5 hours ago

Senior Manager Site Reliability Engineer

GoDaddy 5001-10000 Technology, Information and Internet

GoDaddy is hiring a Senior Manager of Site Reliability Engineering in India to lead its SRE Center of Excellence and build a global, platform-focused operations capability that improves reliability, developer productivity, and scale.

AWS CI/CD Datadog GitHub Actions Go Grafana Kafka Kubernetes Microservices PostgreSQL Prometheus Python Terraform
3 days, 6 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers