qode

qode

qode is a company that focuses on unlocking global opportunities and unleashing potential through no-code solutions. They provide tools and services to help individuals and businesses develop software without the need for traditional coding skills.

Internet Software & Services

Description

  • Design and operate secure, scalable infrastructure for modern applications and AI workloads.
  • Build and maintain automation for CI/CD, infrastructure provisioning, and operational processes.
  • Integrate AI-driven solutions into operational workflows to improve efficiency and delivery speed.
  • Monitor systems, manage incidents, optimize performance, and plan capacity.
  • Establish SRE and DevOps best practices across reproducibility, testing, documentation, and operations.
  • Communicate technical decisions clearly and collaborate cross-functionally to support predictable delivery.
  • Mentor engineers and provide technical leadership to improve platform engineering maturity.

Requirements

  • 6+ years of experience in Site Reliability Engineering, Platform Engineering, or Infrastructure Engineering.
  • Strong preference for an SRE background over a traditional DevOps background.
  • Proven experience designing and operating highly reliable, zero-downtime production systems.
  • Hands-on expertise with Kubernetes or equivalent container orchestration in production.
  • Experience with Infrastructure as Code tools such as Terraform, Pulumi, or CloudFormation.
  • Experience with at least one major cloud provider: AWS, GCP, or Azure.
  • Experience building CI/CD pipelines from the ground up.
  • Experience with observability tools such as Prometheus, Grafana, or Datadog.
  • Hands-on experience supporting or deploying AI/ML workloads such as model inference, vector databases, or GPU workloads.
  • Strong platform security experience, including secrets management, IAM, and runtime hardening.
  • Excellent communication skills and ability to mentor other engineers.
  • Preferred experience with GitOps tools such as ArgoCD or Flux.
  • Preferred true multi-cloud experience across AWS, GCP, and Azure.
  • Preferred experience with multi-cloud API gateways and edge routing.
  • Preferred experience building self-service developer platforms.
  • Preferred familiarity with Node.js, NestJS, or Python for DevOps tooling.
  • Preferred experience collaborating with QA, IT, or ISRM teams on vulnerability remediation and incident investigation.

Benefits

  • Attractive salary range, open to negotiation for strong candidates.
  • Hybrid/remote-friendly work environment.
  • Flexible hours with async teamwork.
  • Work equipment support.
  • Allowance for certification and skill development.
  • Year-end bonus and performance-based rewards.
  • 22 paid leaves from the 5th year, including a full month off.
  • Career growth with personal coaching sessions.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Site Reliability Engineer III

onXmaps 251-1K Food Products

onX is hiring a Site Reliability Engineer to manage the infrastructure, deployment automation, and observability that help developers ship reliably at scale for its outdoor technology products.

Apache Airflow CockroachDB GCP Kubernetes OpenTelemetry Prometheus SQL Terraform
17 hours, 19 minutes ago

Site Reliability Engineer - South Korea

MinIO 51-250 Internet Software & Services

MinIO is hiring a Site Reliability Engineer to help enhance and operate its cloud-native storage platform for high-performance, scalable, and durable data storage and retrieval.

C C++ GitOps Go Kubernetes Microservices Rust
17 hours, 19 minutes ago

Senior Site Reliability Engineer- FedRamp

Veeam Software 1K-5K Internet Software & Services

Veeam is hiring a Site Reliability Engineer to help build its global SRE function for the Veeam Data Cloud, focused on the Government and Sovereign Cloud environment.

Argo CD Azure Bitbucket C# CI/CD ELK Stack Git GitHub Actions GitLab CI GitOps Go Grafana HIPAA Java JavaScript Kubernetes OpenTelemetry Prometheus Pulumi Terraform TypeScript
18 hours, 19 minutes ago

Senior Site Reliability Engineer

RUNWARE 1-10 Internet Software & Services

Runware is hiring a Site Reliability Engineer to keep its AI inference and serverless GPU platform reliable, performant, and resilient as the business scales.

CDN ClickHouse Go Kubernetes Load Balancing Machine Learning MySQL PHP Python RabbitMQ Redis Serverless
18 hours, 49 minutes ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers