Site Reliability Engineer (Remote) - #35039

2 months, 3 weeks ago
Full-time
Mid Level
DevOps and Infrastructure
Recruitment & Search Agency - Headhunter in the Philippines

Recruitment & Search Agency - Headhunter in the Philippines

Manila Recruitment is a top recruitment agency in the Philippines, offering hiring solutions for executive search, IT, developers, managers, and specialized roles. With a database of over 250,000 candidates, we provide innovative headhunting services a...

Professional Services
11-50
Founded 2010

Description

  • Monitor the platform using Cloud Run logs, Temporal workflow UI, GKE pod status, and Pub/Sub queue states.
  • Triage issues to determine whether problems originate in the Python agent layer, Temporal workflows, Go APIs, or Vue frontend.
  • Investigate and resolve paralegal-facing operational issues such as stuck cases, failed faxes, and pending qualifications.
  • Use SQL against AlloyDB PostgreSQL to support troubleshooting and issue investigation.
  • Write and maintain runbooks and escalation procedures for recurring incidents and support workflows.
  • Support integrations across fax, email, SMS/voice, authentication, and external legal or healthcare systems.
  • Work closely with the platform components across backend, workflow, infrastructure, and data services to keep operations running smoothly.

Requirements

  • Experience troubleshooting production systems across logs, workflows, pods, queues, APIs, and UI layers.
  • Comfort working with SQL against PostgreSQL or similar databases.
  • Familiarity with cloud-based infrastructure and services such as GCP, Cloud Run, GKE, Pub/Sub, Redis, and Terraform.
  • Ability to diagnose issues in Python services, Go microservices, and web applications.
  • Experience writing runbooks, support documentation, or escalation procedures.
  • Legal operations, litigation support, or similar domain experience is a bonus.
  • Understanding of integration-based workflows with external systems such as fax, email, SMS/voice, or CRM/CMS tools is preferred.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Site Reliability Engineer 2 (Azure)

PhonePe 5K-10K Capital Markets

PhonePe Limited is hiring a Site Reliability Engineer to manage and scale core cloud infrastructure for a high-volume digital payments environment in India.

Ansible Azure Bash DNS Docker Go Grafana HAProxy InfluxDB Java Linux MySQL Nginx Prometheus Python RabbitMQ SaltStack Terraform Ubuntu
17 hours, 53 minutes ago

Sr. Control System Engineer/Site Reliability Engineer (SRE)

QuEra Computing 11-50 Internet Software & Services

QuEra is seeking a Sr. Control System Engineer/Site Reliability Engineer to integrate and maintain the hardware and software systems that support its quantum control stack and keep development and production environments reliable.

Ansible Bash CI/CD Debian DHCP DNS Docker ELK Stack Embedded Systems Git GitLab CI Go Grafana Jenkins Kubernetes Linux Prometheus Python TCP/IP Terraform Ubuntu
18 hours, 23 minutes ago

Incident Commander

PENN Entertainment 10K-50K Hotels, Restaurants & Leisure

PENN Interactive is hiring an Incident Commander to join its site reliability team and lead cross-functional incident response for its online and physical platforms.

Ansible AWS Docker Elasticsearch GCP Helm JIRA Kafka Kubernetes Linux MySQL PostgreSQL Prometheus Python Redis Terraform
18 hours, 38 minutes ago

Site Reliability Engineer

VantageScore 11-50 Banks

Site Reliability Engineer at a growing engineering team, focused on DevSecOps for maintaining the reliability, security, and compliance of cloud infrastructure, APIs, and software supply chains.

Agile AWS AWS CDK Bash CI/CD CloudFormation CodePipeline Datadog DevSecOps Docker EC2 GitHub Actions Grafana HashiCorp Vault Kong Kubernetes Microservices Python REST API Scrum Terraform
1 day, 18 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers