Jalasoft

Jalasoft

Jalasoft is a world-class technology company with over 20 years of experience in nearshore software development and staff augmentation. They offer Xian Suite, a comprehensive systems management solution, and provide software solutions for small startup...

Internet Software & Services
1K-5K
Founded 2001

Description

  • Ensure the reliability, scalability, and performance of cloud-native platforms running on Microsoft Azure and Kubernetes.
  • Operate Kubernetes workloads in production and troubleshoot workload, resource, and cluster issues.
  • Design and maintain observability with Azure Monitor, Log Analytics, KQL, Prometheus, and Grafana.
  • Define service level indicators, service level objectives, and error budgets.
  • Design alerting and incident response processes, including runbooks and on-call practices.
  • Plan, test, and validate backup, restore, and disaster recovery strategies against RPO and RTO targets.
  • Read, modify, and support infrastructure as code and Azure DevOps pipelines.
  • Automate operational tasks and improvements using scripting.
  • Support incident management and postmortem practices as part of operational excellence.

Requirements

  • 6+ years of experience.
  • 3+ years of experience operating Kubernetes in production.
  • Experience in site reliability engineering or production operations for Kubernetes workloads at scale.
  • Experience with Azure Monitor, Log Analytics, and KQL, including workspace design, data collection rules, and retention strategy.
  • Experience with Prometheus and Grafana, including metrics, exporters, recording rules, alerting rules, and dashboard design.
  • Experience defining and implementing SLIs, SLOs, and error budgets.
  • Experience with alerting and incident response design, including runbooks and on-call practice.
  • Experience with backup, restore, and disaster recovery design and testing, including validation against RPO and RTO targets.
  • Ability to read and modify infrastructure as code using Terraform or Bicep and work with Azure DevOps pipelines.
  • Scripting experience in Python, PowerShell, or Bash.
  • Professional working English.
  • Nice to have: Azure Managed Prometheus and Azure Managed Grafana.
  • Nice to have: OpenTelemetry instrumentation and distributed tracing.
  • Nice to have: Azure Backup, Azure Site Recovery, and snapshot-based VM recovery.
  • Nice to have: Chaos engineering or structured game day practice.
  • Nice to have: Database-layer observability, particularly for Oracle.
  • Nice to have: Cost and capacity management for AKS estates.
  • Nice to have: Incident management tooling and postmortem practice.
  • Nice to have: CKA, AZ-400, or an equivalent certification.

Benefits

  • Remote work.
  • 13 floating holidays.
  • 15 vacation days per year.
  • Good working environment.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Site Reliability Engineer

AssureSoft 51-250 Internet Software & Services

AssureSoft is hiring a remote Site Reliability Engineer to support production cloud infrastructure and platform reliability for long-term client projects.

Argo CD AWS Bash DNS Docker GCP GitHub Actions Go Grafana HTTP Kafka Kubernetes Linux Load Balancing Prometheus Python RabbitMQ Snowflake TCP/IP TLS TypeScript Unix
29 minutes ago

Senior Site Reliability Engineer

Megaport 251-1K Diversified Telecommunication Services

Megaport is hiring a Senior Platform Engineer to support secure, reliable, and maintainable global production systems within its SRE-focused platform team.

AWS Bash Cassandra CI/CD ClickHouse Git GitHub Go Kubernetes Linux PostgreSQL Python Terraform
59 minutes ago

Senior Service Reliability Engineer

Thoughtworks 10K-50K Professional Services

Senior Service Reliability Engineer at Thoughtworks, focused on improving infrastructure reliability, observability, and incident response for customer-facing production systems.

Azure Bash Datadog ELK Stack GitOps Go Grafana Kubernetes Microservices Network Security New Relic Nomad Python REST API Serverless Terraform
1 day, 1 hour ago

Senior Service Reliability Engineer

Thoughtworks 10K-50K Professional Services

Thoughtworks is hiring a Senior Service Reliability Engineer to lead reliability-focused infrastructure work for production systems, improving resilience, observability, incident response, and operational efficiency in support of customer and business goals.

AWS Azure Bash CI/CD Datadog ELK Stack GCP GitOps Go Grafana Kubernetes Microservices New Relic Nomad Python REST API Serverless
1 day, 1 hour ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers