Semios

Semios

Semios is an all-in-one Crop Management Platform that offers a full-service solution for growers to monitor real-time in-field conditions. The platform includes modules for pests, disease, weather, frost, and irrigation, providing insights and data ana...

Food Products
51-250
Founded 2010
$215M raised

Description

  • Lead infrastructure projects and plan higher-risk maintenance work.
  • Contribute to incident response and participate in an on-call rotation.
  • Partner with product and software teams to improve product resiliency and reliability.
  • Mentor team members across SRE practices and technical execution.
  • Analyze production systems to identify reliability, performance, and availability improvements.
  • Drive solutions for components that do not scale.
  • Maintain and improve SLIs aligned to availability and performance targets.
  • Promote automation, refactoring, testing, and small releasable changes to improve quality.
  • Reduce operational overhead through scripting, tooling, and documentation.
  • Manage workload effectively in a remote work environment.

Requirements

  • 8+ years of experience in DevOps, Site Reliability Engineering, or Infrastructure Engineering supporting production cloud environments.
  • 3+ years of experience in a senior or technical leadership role.
  • Hands-on experience with AWS, GCP, or Azure, including deployment, scaling, monitoring, and cost optimization of SaaS applications.
  • Strong knowledge of Linux and bash or similar shell scripting.
  • Strong programming skills in Ruby, Python, Go, or similar languages.
  • Experience with Terraform or similar infrastructure-as-code tools.
  • Experience with Docker, Kubernetes, EKS, or similar container technologies.
  • Experience with CI/CD pipelines such as Buildkite, GitHub Actions, or Jenkins.
  • Strong Git-based version control skills.
  • Experience with observability tools such as Datadog, New Relic, Prometheus, or Splunk and improving SLIs/SLOs.
  • Experience in incident management, on-call rotations, and leading post-incident reviews.
  • Comfort using AI and agentic tooling such as Claude Code to speed up investigation and delivery (nice to have).
  • Hands-on experience with service mesh technologies such as Envoy or Istio (nice to have).
  • Experience with NATS or similar messaging/streaming systems (nice to have).
  • Preferred tech stack includes AWS, GCP/Azure, Terraform, Docker, Kubernetes (EKS), Buildkite/GitHub Actions/Jenkins, Python/Ruby/Go, Datadog/New Relic/Prometheus, GitHub/GitLab, and Linux.

Benefits

  • Salary range of $140,000 to $160,000 per year.
  • Generous vacation policy, company-paid holidays, and a year-end winter break.
  • Hybrid working arrangements with strong work-life balance.
  • Comprehensive health plans covering physical and mental health.
  • Group RRSP with a 3% company match after three months.
  • Supportive, collaborative team environment.
  • Convenient office access via transit and bike paths.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Site Reliability Engineer

PandaDoc 251-1K Internet Software & Services

PandaDoc is hiring a Site Reliability Engineer to help keep its document workflow platform reliable, resilient, and available for customers.

Agile AWS Django Grafana Java Kafka Kubernetes NATS PostgreSQL Python RabbitMQ Spring Boot
16 hours, 42 minutes ago

Senior Site Reliability Engineer

PandaDoc 251-1K Internet Software & Services

PandaDoc is hiring a Site Reliability Engineer to help keep its document workflow platform reliable through incident management, observability, production support, and resilience work across services.

Agile AWS Django Grafana Java Kafka Kubernetes NATS PostgreSQL Python RabbitMQ Spring Boot
16 hours, 42 minutes ago

Senior Site Reliability Engineer

PandaDoc 251-1K Internet Software & Services

PandaDoc is hiring a Site Reliability Engineer to keep its document workflow platform highly available and resilient while supporting production operations and reliability improvements.

Agile AWS Django Grafana Java Kafka Kubernetes NATS PostgreSQL Python RabbitMQ Spring Boot
1 day, 16 hours ago

Senior Site Reliability Engineer

PandaDoc 251-1K Internet Software & Services

PandaDoc is hiring a Site Reliability Engineer to help keep its production document workflow platform reliable, resilient, and low-downtime for customers.

AWS Django Grafana Java Kafka Kubernetes NATS PostgreSQL Python RabbitMQ Spring Boot
1 day, 16 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers