Senior Site Reliability Engineer (DevTools)

3 months, 1 week ago
Full-time
Senior
DevOps and Infrastructure
Nebius

Nebius

Nebius enables B2B companies to build local hyperscaling cloud platforms with cost-effective GPUs, InfiniBand network, and 50% less compute cost. They offer managed Kubernetes and a launch-ready business model for innovative cloud solutions.

Internet Software & Services
51-250

Description

  • Improve services based on user feedback and user problems.
  • Build fault-tolerant, self-healing architecture.
  • Identify ways to speed up systems and reduce user friction.
  • Modify and extend closed-source and open-source solutions, including GitLab and TeamCity plugins.
  • Support users and help resolve their requests and issues.
  • Define metrics that measure user problems and verify that fixes actually resolve them.
  • Work with large-scale build, artifact, and monorepo systems in a production environment.

Requirements

  • Experience combining SRE and software engineering work in roughly a 50/50 split.
  • Experience with Java, Kotlin, Go, Python, and/or Ruby.
  • Understanding of Unix-like systems and the JVM under the hood.
  • Strong focus on improving user experience.
  • Ability to adapt quickly in a fast-changing environment.
  • Experience in Platform Engineering is a plus.
  • Experience operating GitLab or another version control system is a plus.
  • Experience operating TeamCity or another CI system is a plus.
  • Experience with Spring and operating Java monoliths is a plus.
  • Coding interview participation is part of the hiring process.
  • Must be authorized to work in the country of application and provide proof of employment eligibility.

Benefits

  • Competitive compensation.
  • Career growth and learning opportunities.
  • Flexibility and work-life balance.
  • Collaborative and innovative culture.
  • Opportunity to work on impactful AI projects.
  • International environment with talented teams.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Cloud Site Reliability Engineer

Cadwell 51-250 Health Care Providers & Services

Cadwell is seeking a Cloud Site Reliability Engineer to operate and improve AWS infrastructure supporting healthcare customers and ensure reliable, secure, and compliant hosted neurodiagnostic software environments.

AWS Bash CI/CD Encryption HIPAA JavaScript JSON Python Terraform TypeScript YAML
22 hours, 6 minutes ago

Site Reliability Engineer

GiveCampus 51-250 Internet Software & Services

GiveCampus is seeking a hands-on Site Reliability Engineer to strengthen the reliability, performance, observability, and operational maturity of its AWS-based fundraising platform in a remote-first U.S. role.

AWS CI/CD CircleCI Datadog GitHub Actions Kubernetes Linux New Relic OpenSearch PostgreSQL Redis Ruby Ruby on Rails Terraform
22 hours, 36 minutes ago

DevOps / SRE / DevSecOps Engineer (AWS) - Latin America, Remote

Bluelight Consulting 11-50 Internet Software & Services

Bluelight is hiring a DevOps/SRE/DevSecOps professional to operate secure, reliable AWS-based systems and delivery infrastructure for client software projects in a remote consultancy environment.

AWS CI/CD DevSecOps Docker K6 OpenTelemetry PostgreSQL Secrets Management Terraform
1 day, 22 hours ago

DevOps / SRE / DevSecOps Engineer (AWS) - Latin America, Remote

Bluelight Consulting 11-50 Internet Software & Services

Bluelight is hiring a DevOps/SRE/DevSecOps professional to support complex client systems by building secure, reliable, and observable AWS infrastructure and delivery operations.

AWS CI/CD DevSecOps Docker K6 OpenTelemetry PostgreSQL Secrets Management Terraform
1 day, 22 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers