Stellar Cyber

Stellar Cyber

Stellar Cyber provides Next Gen SIEM Security, Network Detection, and Response platforms with AI-driven threat analysis, empowering lean security teams to secure environments effectively.

Professional Services
51-250
Founded 2017
$80M raised

Description

  • Administer and maintain container orchestration platforms and containerized workloads.
  • Monitor, troubleshoot, and support production systems, including participation in on-call rotations.
  • Improve observability across systems and data platforms by enhancing monitoring, logging, and alerting.
  • Administer and optimize cloud environments across multiple providers.
  • Manage and support distributed data platforms and real-time processing systems.
  • Develop and maintain CI/CD pipelines for reliable and efficient deployments.
  • Own and implement Infrastructure as Code practices to improve consistency and scalability.
  • Automate and orchestrate infrastructure using programming and scripting languages.
  • Perform system administration and networking tasks to support internal and external environments.
  • Collaborate with engineers and stakeholders across different time zones to influence architecture, tooling, and best practices.

Requirements

  • 5+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering roles.
  • Proven success operating large-scale production systems in cloud environments such as AWS, GCP, Azure, or OCI.
  • Demonstrated leadership in incident response, on-call best practices, and reliability-focused culture.
  • Strong experience with production on-call operations and incident management.
  • Advanced proficiency in Kubernetes administration and troubleshooting.
  • Hands-on experience with observability tools such as Prometheus, Grafana, Loki, and Alertmanager.
  • Knowledge of chat-based operations interfaces and/or auto-remediation controllers using an AI agentic framework.
  • Understanding of AI agents for auto-triaging alerts, correlating signals, and suggesting root-cause hypotheses.
  • Experience operating data platforms such as Elasticsearch, MongoDB, Spark, Kafka, and Redis.
  • Strong programming and automation skills in Python and Bash.
  • Deep understanding of Infrastructure as Code tools such as Terraform and Helm.
  • Experience with CI/CD pipelines and tooling such as GitHub Actions, Bitbucket, and ArgoCD.
  • Strong technical background in distributed systems, databases, networking, and Linux administration.
  • Excellent problem-solving, communication, and leadership abilities.
  • Bachelor's degree in Computer Science, Engineering, or a related technical field.
  • Certifications in AWS, GCP, observability, Linux, or Kubernetes are a plus.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Site Reliability Specialist II

Everbridge 1K-5K Internet Software & Services

Everbridge is hiring a Senior Site Reliability Engineer II to build resilient platforms and improve the reliability, scalability, and operational readiness of systems supporting critical-event communications.

CI/CD Kubernetes Linux
1 day, 9 hours ago

Site Reliability Engineer (SRE)

Rocket.net 11-50 IT Services

Rocket.net is seeking a Site Reliability Engineer to maintain the reliability and performance of its hosting platform while resolving complex infrastructure issues and providing advanced support to customers.

Apache Bash CDN Cloudflare Datadog DNS Linux MariaDB MySQL Nginx Redis SSH WAF WordPress
1 day, 9 hours ago

Senior Site Reliability Engineer

GoDaddy 5001-10000 Technology, Information and Internet

GoDaddy is hiring a Senior Site Reliability Engineer to support the Commerce ecosystem remotely by improving the reliability, scalability, security, and operation of business-critical production platforms.

Ansible AWS AWS CDK CI/CD CloudFormation Go Kubernetes Linux Pulumi Python SaltStack Terraform TypeScript
2 days, 8 hours ago

Senior Site Reliability Engineer

PENN Entertainment 10K-50K Hotels, Restaurants & Leisure

PENN Entertainment is seeking a Senior Site Reliability Engineer to build and operate the cloud infrastructure supporting large-scale sports betting and media platforms across regulated jurisdictions.

Argo CD AWS Bash CI/CD Datadog GCP GitHub Actions GitOps Go Helm Kubernetes Linux PostgreSQL Python Shell Scripting Terraform
3 days, 8 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers