Sumo Logic

Sumo Logic

Sumo Logic offers top-tier cloud monitoring, log management, and Cloud SIEM tools for web and SaaS apps, empowering businesses with real-time insights and high-quality software delivery.

Internet Software & Services
251-1K
Founded 2010

Description

  • Improve the lifecycle of microservices and related architectural components from design through deployment, operation, and refinement.
  • Define, evolve, and manage service level objectives (SLOs).
  • Write code and automation to reduce operational workload, improve efficiency, strengthen security posture, and eliminate toil.
  • Scale systems sustainably through automation and reliability-focused improvements.
  • Facilitate blame-free root cause analysis meetings and drive learning from incidents.
  • Participate in and improve global incident response coordination across products.
  • Drive root cause identification and issue resolution with cross-functional teams.
  • Work closely with multiple teams to optimize the operations of their microservices.
  • Operate in a fast-paced, iterative environment.

Requirements

  • 6+ years of industry experience.
  • Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or another scientific or technical discipline.
  • Cloud-native application development experience using best practices and design patterns.
  • Strong debugging and troubleshooting skills across the full technology stack.
  • Deep understanding of AWS networking, compute, storage, and managed services.
  • Experience with modern CI/CD tooling such as Kubernetes, Terraform, Ansible, and Jenkins.
  • Experience with full lifecycle support of services, from creation to production support.
  • Infrastructure as Code experience with tools such as Terraform or AWS CloudFormation.
  • Ability to author production-ready code in at least one of Java, Scala, or Go.
  • Experience with Linux systems and command-line work.
  • Understanding of modern cloud-native software security practices.
  • Experience with agile frameworks such as Scrum and Kanban.
  • Flexibility to step into new roles and responsibilities.
  • Willingness to learn and use Sumo Logic products to solve reliability and security issues.
  • Preferred: experience using Sumo Logic or other observability products for reliability and security.
  • Preferred: experience with planet-scale product development.
  • Preferred: expert-level experience running and operating SaaS products on AWS.
  • Preferred: experience with streaming technologies such as Kafka, Kafka Streams, or KSQL.
  • Preferred: expert-level experience in one or more of Java, Go, Scala, or Python.
  • Preferred: expert-level experience in one or more of Terraform, Jenkins, or Kubernetes.
  • Preferred: extensive experience running and tuning JVM workloads at scale.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Site Reliability Engineer

Counterpart Health 51-200 hospital & health care

Counterpart Health is hiring a Senior Site Reliability and Infrastructure Engineer to support and evolve the technology platform behind its primary care tool and maintain reliable infrastructure for domestic and international workloads.

AWS Azure CI/CD Containerd DNS Docker GCP Go gRPC Helm Kubernetes Linux Load Balancing Prometheus Python Shell Scripting TCP/IP
15 hours, 31 minutes ago

Senior Test Platform & Reliability Engineer - Star Trek Fleet Command

Scopely 1K-5K Internet Software & Services

Scopely is hiring a Senior Test Platform & Reliability Engineer in Ireland to build validation, reliability, and developer enablement platforms for Star Trek Fleet Command’s large-scale live-service backend systems.

AWS Bash CI/CD Docker GitLab Go Python Terraform
15 hours, 46 minutes ago

Senior Software Engineer - Databases, SRE | Canada | Remote

Grafana 1K-5K IT Services

Grafana Labs is hiring a Senior Software Engineer for its remote SRE team to improve reliability and operability of Grafana Cloud database services for high-SLA customers across AWS, GCP, and Azure.

AWS Azure GCP Go Helm Java Kubernetes Linux Microservices Python Terraform
1 day, 14 hours ago

Senior Site Reliability Engineer

Semios 51-250 Food Products

Semios Group is hiring a Senior Site Reliability Engineer to help scale, secure, and improve the reliability of its global agricultural technology platform.

AWS Azure Bash Buildkite CI/CD Datadog Docker Envoy GCP Git GitHub GitHub Actions GitLab Go Jenkins Kubernetes Linux NATS New Relic Prometheus Python Ruby Splunk Terraform
1 day, 16 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers