Staff Site Reliability Engineer

10 hours, 12 minutes ago
Full-time
Lead
DevOps and Infrastructure
Caseware

Caseware

CaseWare International Inc. provides cutting-edge software solutions for accounting firms, corporations, and governments, enabling users worldwide to work smarter and transform insights into impact.

Internet Software & Services
251-1K
Founded 1988

Description

  • Drive reliability engineering and operational excellence for mission-critical AWS and Kubernetes services.
  • Design and improve deployment, release, and rollback strategies for distributed systems.
  • Build secure-by-default CI/CD pipelines with automation, governance, and policy controls.
  • Improve observability through metrics, logs, tracing, and actionable alerting.
  • Define and mature SLIs, SLOs, reliability standards, runbooks, and reliability metrics.
  • Lead high-severity incident response, communication, resolution, and post-incident reviews.
  • Partner with Engineering, Security, Platform, and Product teams to improve resilience, performance, and platform standards.
  • Mentor engineers on cloud-native technologies, SRE principles, and operational practices.
  • Build scalable backend services, APIs, event-driven systems, and internal platform tooling.

Requirements

  • 8+ years of experience in SRE, Platform Engineering, DevOps, or related cloud-native roles.
  • Deep AWS expertise, including EKS, IAM, VPC, Lambda, CloudFront, S3, networking, and security.
  • Advanced experience operating and scaling production Kubernetes environments.
  • Strong hands-on experience with Istio service mesh, including traffic management, security, observability, and resilience.
  • Experience with Infrastructure as Code, preferably AWS CDK.
  • Experience building CI/CD pipelines with GitHub Actions or similar tools.
  • Strong TypeScript and Node.js proficiency for platform engineering and automation.
  • Experience with CloudWatch, OpenTelemetry, AWS X-Ray, and Kubernetes observability tooling.
  • Knowledge of zero-trust architectures, mTLS, Kubernetes networking, resilience engineering, and distributed-system troubleshooting.
  • Preferred: progressive delivery, regulated SaaS, FinOps, internal developer platforms, or certifications such as CKA, CKAD, CKS, KCSA, KCNA, or Kubestronaut.

Benefits

  • Annual base salary of $140,000–$155,000 CAD.
  • Eligibility for discretionary bonus and/or commission.
  • Comprehensive health insurance and retirement plans.
  • Remote work and flexible work options with generous time off.
  • Career growth, recognition programs, and performance bonuses.
  • Collaborative, inclusive culture with global project opportunities.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Sr Lead Network Reliability Engineer

Coupa Software 1K-5K Internet Software & Services

Coupa is hiring a Sr. Lead Network Development Engineer to scale and operate its global SaaS platform’s cloud networking infrastructure through automation, reliability engineering, and technical leadership.

Ansible AWS Azure Chef DNS Fortinet Go Java Kubernetes Linux Python Ruby TCP/IP Terraform TLS
9 hours, 57 minutes ago

Senior Monitoring/Observability Architect

Makpar 51-250 Internet Software & Services

Makpar is seeking a Senior Monitoring/Observability Architect to lead enterprise monitoring strategy, architecture, and implementation guidance for a large federal government program.

Datadog Splunk
9 hours, 57 minutes ago

Senior Parts and Reliability Engineer

SCOUT 11-50 Aerospace & Defense

Scout Space is seeking a Parts and Reliability Engineering Subject Matter Expert to establish and lead mission-assurance processes for electronic, electromechanical, and mechanical parts used in optical payloads and launch-vehicle components.

1 day, 10 hours ago

Senior Site Reliability Engineer

Lodgify 251-1K Internet Software & Services

Lodgify, a Barcelona-based vacation-rental technology company, is hiring a Senior Site Reliability Engineer to improve the reliability, scalability, observability, and operational ownership of its cloud platform and critical product services.

Datadog Grafana Kubernetes Microservices Prometheus Python
2 days, 10 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers