Thoughtworks

Thoughtworks

Thoughtworks is a leading technology consultancy that focuses on transforming digital journeys for clients by delivering innovative software design, engineering excellence, and strategic insights to address pressing global challenges.

Professional Services
10K-50K
Founded 1993
$748M raised

Description

  • Improve site reliability by designing fault-tolerant mechanisms and architectures that reduce mean time to detect and mean time to respond.
  • Drive the integration of observability automation into the CI/CD pipeline.
  • Handle production incidents, manage client communications during incidents, and draft root cause analysis documents.
  • Monitor production system performance and improve scaling to meet SLA and SLO targets.
  • Advise application development teams on reliability improvements and help implement them.
  • Improve observability through logging, metrics, and alert tuning to reduce false alarms and operational toil.
  • Implement and operationalize chaos engineering practices to test system reliability regularly.
  • Set direction for site reliability in line with client goals, including high-availability targets where required.

Requirements

  • Hands-on experience with programming and scripting languages such as Python, Go, or Bash.
  • Good understanding of at least one public cloud platform such as AWS, Azure, or GCP.
  • Experience with observability tools such as Grafana, Datadog, New Relic, ELK Stack, or Dynatrace.
  • Familiarity with DevOps and GitOps practices.
  • Strong knowledge of container-based architecture and orchestration tools such as Kubernetes, AWS EKS, Docker Swarm, or Nomad.
  • Understanding of modern architecture and design patterns, including microservices, serverless functions, NoSQL, and RESTful APIs.
  • Experience with infrastructure aligned to cloud well-architected principles, including reliability, security, cost optimization, performance efficiency, and operations.
  • Strong communication and articulation skills, with proficiency in English.
  • Ability to collaborate across multiple cross-functional teams and negotiate effectively.
  • Ability to work under pressure during production incidents and maintain composure.
  • Willingness to be part of a rotation-based, need-based 24x7 available team.

Benefits

  • Career development supported by interactive tools, development programs, and teammates who help you grow.
  • Flexible career path with autonomy over how you develop your role.
  • Inclusive, supportive, and collaborative culture.
  • Opportunity to work for a leading technology consultancy on impactful client problems.
  • Remote work indicated by the #LI-Remote tag.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Site Reliability Engineer

AssureSoft 51-250 Internet Software & Services

AssureSoft is hiring a remote Site Reliability Engineer to support production cloud infrastructure and platform reliability for long-term client projects.

Argo CD AWS Bash DNS Docker GCP GitHub Actions Go Grafana HTTP Kafka Kubernetes Linux Load Balancing Prometheus Python RabbitMQ Snowflake TCP/IP TLS TypeScript Unix
29 minutes ago

Senior Site Reliability Engineer

Megaport 251-1K Diversified Telecommunication Services

Megaport is hiring a Senior Platform Engineer to support secure, reliable, and maintainable global production systems within its SRE-focused platform team.

AWS Bash Cassandra CI/CD ClickHouse Git GitHub Go Kubernetes Linux PostgreSQL Python Terraform
59 minutes ago

Site Reliability Engineer - Azure, Observability and Scripting

Jalasoft 1K-5K Internet Software & Services

Jalasoft is hiring a Site Reliability Engineer to support the reliability, scalability, and performance of cloud-native platforms on Microsoft Azure and Kubernetes.

Azure Bash Grafana Kubernetes OpenTelemetry Oracle PowerShell Prometheus Python Terraform
1 day ago

Senior Site Reliability Engineer

Aspenview Technology Partners Internet Software & Services

AspenView Technology Partners is hiring a Senior Site Reliability Engineer to support a large-scale cloud transformation by building and operating resilient, secure platforms for critical enterprise applications.

Agile AWS Bash Datadog DevSecOps GCP Grafana Kubernetes Prometheus Python Splunk Terraform
2 days ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers