AWS - Incident Handler

2 days, 9 hours ago
Full-time
Senior
DevOps and Infrastructure
Caseware

Caseware

CaseWare International Inc. provides cutting-edge software solutions for accounting firms, corporations, and governments, enabling users worldwide to work smarter and transform insights into impact.

Internet Software & Services
251-1K
Founded 1988

Description

  • Initiate and oversee incident response as the primary coordinator in a 24/7 SaaS environment.
  • Coordinate engineers, product managers, support teams, and other stakeholders during incidents.
  • Use and integrate tools including JIRA, PagerDuty, New Relic, AWS, and Microsoft Teams to manage incident response.
  • Drive teams toward rapid and efficient incident resolution.
  • Guide resolution strategies using knowledge of the software and infrastructure environment.
  • Communicate incident status, resolution plans, and updates to internal and external stakeholders.
  • Track and report uptime, reliability, and performance metrics.
  • Lead post-mortems and document root causes, timelines, impacts, lessons learned, and remediation actions.
  • Follow up on corrective actions and implement proactive measures to reduce risk and improve system resilience.

Requirements

  • 5+ years of experience managing critical incidents in SaaS environments.
  • Experience with cloud environments such as AWS, DevOps practices, or technical operations.
  • Experience in a similar incident management role, preferably in a software or technology company.
  • Strong technical background in incident management and response.
  • Proven ability to lead teams through rapid incident resolution.
  • Understanding of modern software environments and JIRA and PagerDuty integrations.
  • Excellent written and verbal communication skills.
  • Ability to remain effective under pressure while managing competing priorities.
  • Successful candidates must complete a background check, typically including identity verification and a criminal record check.

Benefits

  • Indefinite-term contract with legal benefits, prepaid medical coverage, life insurance, and funeral assistance.
  • Fully remote work in Colombia with internet and home-office allowances.
  • Competitive compensation above the market average.
  • Five personal days annually, paid sick-leave top-up, service recognition time off, and increased vacation after five years.
  • Training budget, mentorship, career growth, and recognition programs.
  • AI-first engineering environment using modern tools, automation, and AI-driven workflows.
  • Collaborative, inclusive culture with flexible work options and international projects.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Site Reliability Engineer

Lodgify 251-1K Internet Software & Services

Lodgify, a Barcelona-based vacation-rental technology company, is hiring a Senior Site Reliability Engineer to improve the reliability, scalability, observability, and operational ownership of its cloud platform and critical product services.

Datadog Grafana Kubernetes Microservices Prometheus Python
10 hours, 7 minutes ago

Senior Site Reliability Engineer

PointClickCare 1K-5K Health Care Providers & Services

PointClickCare is seeking a Senior Site Reliability Engineer to provide technical leadership and improve the reliability, automation, observability, and operational efficiency of cloud-based healthcare applications.

Agile Ansible AWS Azure C C++ Chef Docker Go Java Kubernetes Linux Perl Puppet Python Ruby TCP/IP Terraform Windows Server
10 hours, 22 minutes ago

Principal Site Reliability Engineer, Platform

Blue River Technology 251-1K Industrial Conglomerates

Blue River Technology, a John Deere company developing AI and robotics for agriculture and construction, is seeking a Principal Site Reliability Engineer to shape and scale the Kubernetes-based platform that enables reliable delivery of autonomous systems and products.

Argo CD AWS CI/CD GitHub Actions Go JavaScript Kubernetes Python Rust Terraform
2 days, 10 hours ago

Senior Site Reliability Engineer - Cloud Platform

GoDaddy 5001-10000 Technology, Information and Internet

GoDaddy’s Global Compute team is seeking a remote infrastructure engineer to operate and scale AWS production infrastructure that supports the company’s engineering teams.

AWS AWS CDK CI/CD CloudFormation GitOps Python
2 days, 10 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers