Geotab

Geotab

Geotab is a leading provider of GPS fleet tracking and management solutions, leveraging data analytics and machine learning to optimize fleet performance, enhance driver safety, and ensure regulatory compliance worldwide.

Road & Rail
1K-5K
Founded 2000

Description

  • Define and own the enterprise-wide observability architecture, including technical standards, reference architectures, and multi-year roadmaps.
  • Evaluate, select, and standardize observability tools to reduce tool sprawl and optimize total cost of ownership.
  • Design scalable data pipelines and storage strategies for ingesting and querying petabyte-scale telemetry data across metrics, traces, logs, and profiling.
  • Design Terraform modules and Helm charts for declarative observability infrastructure provisioning across multi-cloud environments.
  • Establish and enforce instrumentation standards using OpenTelemetry, including SDK guidelines, collector deployment patterns, and semantic conventions.
  • Define and champion SLO, SLI, and error-budget frameworks across engineering teams.
  • Serve as a senior escalation point during critical incidents to accelerate diagnosis and resolution.
  • Provide architectural mentorship and technical guidance to Observability Engineers and SRE team members.
  • Collaborate closely with SRE, platform engineering, application development, security, and compliance stakeholders.
  • Influence and drive technical direction across multiple teams and organizational boundaries.

Requirements

  • 5-8 years of experience in Observability Architecture, Site Reliability Engineering, or Platform/Infrastructure Engineering.
  • Post-secondary diploma or degree in Engineering, Computer Science, or a related field.
  • Mastery of the OpenTelemetry ecosystem and expert-level knowledge of Prometheus-compatible metrics systems such as VictoriaMetrics and Thanos.
  • Advanced experience with tracing systems such as Grafana Tempo and Jaeger, and log aggregation platforms such as Loki, Elasticsearch, and Google BigQuery.
  • Expert-level proficiency in cloud infrastructure, with GCP strongly preferred, and Kubernetes architecture.
  • Strong software engineering skills in Go, Python, or similar languages for building cloud-native tooling.
  • Excellent communication skills with the ability to articulate technical architecture to executive audiences and influence across organizational boundaries.
  • Deep expertise in designing enterprise-scale observability platforms.
  • Preferred certifications: Google Cloud Professional Cloud Architect or Certified Kubernetes Administrator (CKA).
  • Ability to work in a fast-paced, evolving environment with willingness to take on new tasks and activities.

Benefits

  • Hiring range of $116,200 to $155,000 CAD annually.
  • Flex working arrangements and a flexible hybrid working model.
  • Home office reimbursement program.
  • Baby bonus and parental leave top-up program.
  • Online learning and networking opportunities.
  • Electric vehicle purchase incentive program.
  • Competitive medical and dental benefits.
  • Retirement savings program.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Ssr Monitoring and Observability Analyst

Coderio 51-250 Internet Software & Services

Coderio is hiring an Observability & Monitoring Analyst to design and operate monitoring systems that improve availability, performance, and incident response across global clients’ IT environments.

AWS Azure Bash Datadog DNS Docker ELK Stack Fluentd GCP Grafana Jaeger Kibana Kubernetes Linux Load Balancing Logstash New Relic OpenTelemetry Prometheus Python TCP/IP Zipkin
49 minutes ago

Senior Site Reliability Engineer

Counterpart Health 51-200 hospital & health care

Counterpart Health is hiring a Senior Site Reliability and Infrastructure Engineer to support and evolve the technology platform behind its primary care tool and maintain reliable infrastructure for domestic and international workloads.

AWS Azure CI/CD Containerd DNS Docker GCP Go gRPC Helm Kubernetes Linux Load Balancing Prometheus Python Shell Scripting TCP/IP
2 days ago

Senior Test Platform & Reliability Engineer - Star Trek Fleet Command

Scopely 1K-5K Internet Software & Services

Scopely is hiring a Senior Test Platform & Reliability Engineer in Ireland to build validation, reliability, and developer enablement platforms for Star Trek Fleet Command’s large-scale live-service backend systems.

AWS Bash CI/CD Docker GitLab Go Python Terraform
2 days ago

Senior Software Engineer - Databases, SRE | Canada | Remote

Grafana 1K-5K IT Services

Grafana Labs is hiring a Senior Software Engineer for its remote SRE team to improve reliability and operability of Grafana Cloud database services for high-SLA customers across AWS, GCP, and Azure.

AWS Azure GCP Go Helm Java Kubernetes Linux Microservices Python Terraform
2 days, 23 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers