Jalasoft

Jalasoft

Jalasoft is a world-class technology company with over 20 years of experience in nearshore software development and staff augmentation. They offer Xian Suite, a comprehensive systems management solution, and provide software solutions for small startup...

Internet Software & Services
1K-5K
Founded 2001

Description

  • Ensure the reliability, scalability, and performance of cloud-native platforms running on Microsoft Azure and Kubernetes.
  • Operate Kubernetes workloads in production and troubleshoot workload, resource, and cluster issues.
  • Design and maintain observability with Azure Monitor, Log Analytics, KQL, Prometheus, and Grafana.
  • Define service level indicators, service level objectives, and error budgets.
  • Design alerting and incident response processes, including runbooks and on-call practices.
  • Plan, test, and validate backup, restore, and disaster recovery strategies against RPO and RTO targets.
  • Read, modify, and support infrastructure as code and Azure DevOps pipelines.
  • Automate operational tasks and improvements using scripting.
  • Support incident management and postmortem practices as part of operational excellence.

Requirements

  • 6+ years of experience.
  • 3+ years of experience operating Kubernetes in production.
  • Experience in site reliability engineering or production operations for Kubernetes workloads at scale.
  • Experience with Azure Monitor, Log Analytics, and KQL, including workspace design, data collection rules, and retention strategy.
  • Experience with Prometheus and Grafana, including metrics, exporters, recording rules, alerting rules, and dashboard design.
  • Experience defining and implementing SLIs, SLOs, and error budgets.
  • Experience with alerting and incident response design, including runbooks and on-call practice.
  • Experience with backup, restore, and disaster recovery design and testing, including validation against RPO and RTO targets.
  • Ability to read and modify infrastructure as code using Terraform or Bicep and work with Azure DevOps pipelines.
  • Scripting experience in Python, PowerShell, or Bash.
  • Professional working English.
  • Nice to have: Azure Managed Prometheus and Azure Managed Grafana.
  • Nice to have: OpenTelemetry instrumentation and distributed tracing.
  • Nice to have: Azure Backup, Azure Site Recovery, and snapshot-based VM recovery.
  • Nice to have: Chaos engineering or structured game day practice.
  • Nice to have: Database-layer observability, particularly for Oracle.
  • Nice to have: Cost and capacity management for AKS estates.
  • Nice to have: Incident management tooling and postmortem practice.
  • Nice to have: CKA, AZ-400, or an equivalent certification.

Benefits

  • Remote work.
  • 13 floating holidays.
  • 15 vacation days per year.
  • Good working environment.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Cloud Site Reliability Engineer

Cadwell 51-250 Health Care Providers & Services

Cadwell is seeking a Cloud Site Reliability Engineer to operate and improve AWS infrastructure supporting healthcare customers and ensure reliable, secure, and compliant hosted neurodiagnostic software environments.

AWS Bash CI/CD Encryption HIPAA JavaScript JSON Python Terraform TypeScript YAML
7 minutes ago

Site Reliability Engineer

GiveCampus 51-250 Internet Software & Services

GiveCampus is seeking a hands-on Site Reliability Engineer to strengthen the reliability, performance, observability, and operational maturity of its AWS-based fundraising platform in a remote-first U.S. role.

AWS CI/CD CircleCI Datadog GitHub Actions Kubernetes Linux New Relic OpenSearch PostgreSQL Redis Ruby Ruby on Rails Terraform
37 minutes ago

DevOps / SRE / DevSecOps Engineer (AWS) - Latin America, Remote

Bluelight Consulting 11-50 Internet Software & Services

Bluelight is hiring a DevOps/SRE/DevSecOps professional to operate secure, reliable AWS-based systems and delivery infrastructure for client software projects in a remote consultancy environment.

AWS CI/CD DevSecOps Docker K6 OpenTelemetry PostgreSQL Secrets Management Terraform
1 day ago

DevOps / SRE / DevSecOps Engineer (AWS) - Latin America, Remote

Bluelight Consulting 11-50 Internet Software & Services

Bluelight is hiring a DevOps/SRE/DevSecOps professional to support complex client systems by building secure, reliable, and observable AWS infrastructure and delivery operations.

AWS CI/CD DevSecOps Docker K6 OpenTelemetry PostgreSQL Secrets Management Terraform
1 day ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers