Thoughtworks

Thoughtworks

Thoughtworks is a leading technology consultancy that focuses on transforming digital journeys for clients by delivering innovative software design, engineering excellence, and strategic insights to address pressing global challenges.

Professional Services
10K-50K
Founded 1993
$748M raised

Description

  • Understand SRE goals and requirements from both technical and business perspectives.
  • Identify and implement reliability improvements, including fault-tolerant mechanisms and architectures.
  • Improve incident management through prioritization, triage, communication, mitigation, post-mortems, and corrective actions.
  • Manage client stakeholder expectations during production incidents and provide technical analysis and remediation plans.
  • Act as an interface for senior client stakeholders and C-level executives when needed.
  • Liaise with client engineering teams and build trust with senior stakeholders and team leads.
  • Identify opportunities to improve system performance and reliability against SLAs, SLOs, KPIs, and business objectives.
  • Collaborate with application development leads and solution architects on reliability-focused design changes and best practices.
  • Guide and assist SRE teams in implementing improvements.
  • Mentor and support the growth of other SREs on the team.

Requirements

  • Programming experience in one or more high-level languages such as Python, Golang, Shell scripting, Ruby, or Java.
  • Familiarity with DevOps and GitOps practices, including observability automation in CI/CD tools such as GitLab, Jenkins, or CircleCI.
  • In-depth knowledge of configuration management and Infrastructure as Code tools such as Terraform, Ansible, ARM, and CloudFormation.
  • Expertise with observability, logging, tracing, and monitoring tools such as Grafana, Prometheus, Graylog, Jaeger, Zipkin, or the ELK stack.
  • Hands-on experience with container-based architecture and orchestration tools such as Kubernetes, AWS EKS, Docker Swarm, or Nomad.
  • Experience tuning and scaling applications and infrastructure for heavy-load scenarios, including periodic traffic spikes and tsunami patterns.
  • Understanding of SLI/SLO/SLA quality gates, chaos engineering, golden signals, blameless postmortems, synthetic monitoring, distributed tracing, end-user monitoring, and performance testing.
  • Experience with network load balancing, security stacks, TLS, certificate management, and standard networking protocols.
  • Strong communication, articulation, listening, and presentation skills with proficiency in English.
  • Ability to work under pressure during production incidents and participate in a rotation-based, 24x7 available team.

Benefits

  • Annual salary range of $155,000 to $249,000 USD.
  • Learning and development support with interactive tools and development programs.
  • Remote work indicated by the listing.
  • Reasonable accommodations for qualified applicants with disabilities or sincerely held religious beliefs.
  • Equal-opportunity employer committed to inclusive hiring and a discrimination-free workplace.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Site Reliability Engineer

Coforge 10K-50K IT Services

Coforge is hiring a remote Site Reliability Engineer to help build and operate reliable cloud platforms and production systems for teams across Costa Rica, Peru, Colombia, and Bolivia.

AWS Azure C# CI/CD Datadog GCP Go Grafana Java Kubernetes Linux Microservices New Relic OpenTelemetry Prometheus Python Splunk Terraform TypeScript
5 hours, 5 minutes ago

Sr. Staff Platform/Data Reliability Engineer, Databricks (R5537)

Bitly 51-250 Internet Software & Services

Shield AI is seeking a Sr. Staff Platform / Data Reliability Engineer to make its Databricks platform reliable, secure, scalable, and operationally mature for enterprise and regulated use.

CI/CD Databricks Git
2 days, 5 hours ago

Senior Manager Site Reliability Engineer

GoDaddy 5001-10000 Technology, Information and Internet

GoDaddy is hiring a Senior Manager of Site Reliability Engineering in India to lead its SRE Center of Excellence and build a global, platform-focused operations capability that improves reliability, developer productivity, and scale.

AWS CI/CD Datadog GitHub Actions Go Grafana Kafka Kubernetes Microservices PostgreSQL Prometheus Python Terraform
3 days, 5 hours ago

Vice President, Global Production Operations & Reliability

Everbridge 1K-5K Internet Software & Services

Everbridge is hiring a Vice President, Global Production Operations & Reliability to lead the company’s global production operations for its cloud-native SaaS platform and drive reliability, scalability, security, and operational excellence.

AWS CI/CD Kubernetes
4 days, 4 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers