Site Reliability Engineer - Canada Wide - Remote

4 months, 4 weeks ago
Senior
DevOps and Infrastructure
Newton

Newton

Newton provides a user-friendly platform for Canadians to buy and sell Bitcoin, Ethereum, and over 70 other cryptocurrencies, offering competitive trading fees and a seamless trading experience.

Capital Markets
51-250
Founded 2018
$35M raised

Description

  • Implement improvements to infrastructure reliability, fault tolerance, scalability, and performance.
  • Manage incidents and coordinate the appropriate teams during operational issues.
  • Respond to automated alerts and support critical services through an on-call rotation.
  • Define and maintain SLIs, SLOs, SLAs, and error budgets to guide reliability decisions.
  • Improve observability across systems through metrics, logs, and tracing.
  • Reduce production issue detection, troubleshooting, and resolution time.
  • Enhance monitoring, alerting, dashboards, tracing, and runbooks for critical services.
  • Lead postmortems and follow-up actions to prevent repeat incidents.
  • Automate manual operational practices and help build better incident response processes.
  • Work closely with engineering teams to improve system design and operational excellence.

Requirements

  • Experience designing and operating scalable, reliable systems in AWS or a similar cloud environment.
  • Experience handling on-call shifts for critical systems.
  • Experience with chaos engineering tools or practices, such as Gremlin.
  • Ability to debug live production systems.
  • Experience writing and deploying code with zero downtime.
  • Experience scripting or developing with Linux Shell, Python, JavaScript, Java, or similar languages.
  • Self-starter with the ability to take initiative in ambiguous environments, preferably in a startup setting.

Benefits

  • Remote work across Canada.
  • Inclusive work environment that welcomes candidates from all backgrounds and perspectives.
  • Reasonable accommodations provided during the application process for candidates who need assistance.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

SRE [Antifraud]

Banco Plata, S.A., Institución de Banca Múltiple. 1001-5000 Banking / financial services

Join the Antifraud team as an SRE, helping operate scalable AML, Anti-Fraud, and KYC platforms that perform real-time financial crime prevention checks without slowing transactions.

Bash ELK Stack Helm Kafka Kubernetes Load Balancing Prometheus Python Secrets Management
22 hours, 29 minutes ago

Staff Site Reliability Engineer

Zscaler 1K-5K Internet Software & Services

Zscaler is seeking a remote Staff Site Reliability Engineer in the Netherlands to build, secure, automate, and operate scalable Linux, Kubernetes, and cloud infrastructure for its global security platform.

Ansible Bash DHCP Docker Go HashiCorp Vault Kubernetes Linux Python Secrets Management SSH
1 day, 23 hours ago

Staff Site Reliability Engineer-Observability

GoDaddy 5001-10000 Technology, Information and Internet

GoDaddy is hiring a remote Reliability Engineer in India to operate and modernize the monitoring, compliance, infrastructure, and incident-response systems supporting its global Domains platform.

Ansible Argo CD AWS AWS CDK CI/CD Cybersecurity Elasticsearch Git Go Gradle Grafana Jenkins Kubernetes Maven MySQL PostgreSQL Prometheus Python Ruby SQL Terraform
2 days, 23 hours ago

Site Reliability Engineer - AWS and Azure

Jalasoft 1K-5K Internet Software & Services

Jalasoft is hiring a Site Reliability Engineer to build, maintain, and improve reliable, scalable, and secure cloud infrastructure across Windows and Linux environments.

Ansible AWS Azure Bash CI/CD Docker GitHub GitHub Actions Grafana Kubernetes PowerShell Prometheus Python TeamCity Terraform
3 days, 23 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers