Nebius

Nebius

Nebius enables B2B companies to build local hyperscaling cloud platforms with cost-effective GPUs, InfiniBand network, and 50% less compute cost. They offer managed Kubernetes and a launch-ready business model for innovative cloud solutions.

Internet Software & Services
51-250

Description

  • Improve services based on user feedback and user problems.
  • Build fault-tolerant, self-healing architecture.
  • Identify ways to speed up systems and reduce user friction.
  • Modify and extend closed-source and open-source solutions, including GitLab and TeamCity plugins.
  • Support users and help resolve their requests and issues.
  • Define metrics that measure user problems and verify that fixes actually resolve them.
  • Work with large-scale build, artifact, and monorepo systems in a production environment.

Requirements

  • Experience combining SRE and software engineering work in roughly a 50/50 split.
  • Experience with Java, Kotlin, Go, Python, and/or Ruby.
  • Understanding of Unix-like systems and the JVM under the hood.
  • Strong focus on improving user experience.
  • Ability to adapt quickly in a fast-changing environment.
  • Experience in Platform Engineering is a plus.
  • Experience operating GitLab or another version control system is a plus.
  • Experience operating TeamCity or another CI system is a plus.
  • Experience with Spring and operating Java monoliths is a plus.
  • Coding interview participation is part of the hiring process.
  • Must be authorized to work in the country of application and provide proof of employment eligibility.

Benefits

  • Competitive compensation.
  • Career growth and learning opportunities.
  • Flexibility and work-life balance.
  • Collaborative and innovative culture.
  • Opportunity to work on impactful AI projects.
  • International environment with talented teams.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

AI infrastructure Engineer (SRE) Bangalore

Together 1-10 IT Services

Together AI is hiring an AI Infrastructure Engineer (SRE) to keep its user-facing services and production systems reliable, scalable, and available as the company builds next-generation AI infrastructure.

Ansible Kubernetes Machine Learning PagerDuty Terraform
6 hours, 56 minutes ago

Senior Site Reliability Engineer

Omilia 251-1K IT Services

Omilia is hiring a Senior Site Reliability Engineer to operate and improve cloud-based production platforms, observability, and reliability practices across development and engineering teams.

Agile Ansible AWS Bash CentOS Go Grafana Kubernetes MySQL PostgreSQL Prometheus Python Redis TCP/IP Terraform
7 hours, 26 minutes ago

Site Reliability Engineer II

MRSOOL 1K-5K Air Freight & Logistics

Mrsool is hiring an experienced Site Reliability Engineer to help ensure the stability and reliability of its delivery platform while supporting feature delivery and infrastructure growth.

Ansible AWS Azure Chef Docker GCP Go Grafana Java Kubernetes Nagios Prometheus Puppet Python Ruby Terraform
7 hours, 56 minutes ago

Senior Site Reliability Engineer

Latitude AI 501-1000 information technology & services

Latitude AI is hiring a Site Reliability Engineer to help operate and improve the mission-critical systems behind Ford’s autonomous driving platform.

AWS CloudFormation Elasticsearch GCP Go Jaeger Kubernetes Linux Prometheus Python TCP/IP Terraform
1 day, 6 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers