Ensono

Ensono

Ensono provides comprehensive hybrid IT solutions and governance, enabling businesses to navigate complexity and modernize their technology infrastructure, from cloud services to mainframe systems, tailored to each client's unique journey.

IT Services
1K-5K
Founded 1969

Description

  • Engineer and operate a scalable monitoring and observability platform for Ensono’s Hybrid Cloud clients.
  • Plan and execute the strategic roadmap for observability and monitoring tools in alignment with business and client requirements.
  • Define monitoring best practices, including proactive alerting, anomaly detection, and performance analytics.
  • Operate and optimize end-to-end monitoring solutions for real-time visibility into network, distributed systems, and applications.
  • Establish automated alerting thresholds based on Service Level Objectives (SLOs) and Service Level Agreements (SLAs).
  • Establish monitoring audit standards for conformance and compliance across standard and custom monitors.
  • Serve as the point of escalation for day-to-day monitoring-related incidents.
  • Automate monitoring configurations and telemetry collection using scripting and Infrastructure as Code tools such as Ansible and Terraform.

Requirements

  • 7+ years of experience in observability or monitoring engineering operational roles.
  • 7+ years of hands-on experience with ITSM platforms such as ServiceNow and monitoring tools such as BMC, Data Dog, Entuity, or similar tools.
  • Strong proficiency in Python, Bash, and JavaScript for automation and scripting.
  • Experience with Infrastructure as Code tools such as Ansible and Terraform for observability tool deployment.
  • Strong analytical and problem-solving skills for diagnosing complex issues.
  • Effective communication and leadership skills, especially for training and cross-functional collaboration.
  • Ability to think holistically about business-impacting processes and continuously refine them.
  • Ability to thrive in an independent and collaborative fast-paced environment while managing priorities effectively.
  • Bachelor’s degree in a related field.
  • Master’s degree in an information technology-related field (preferred).
  • Proficiency in cloud platforms such as AWS, Azure, or GCP, and Kubernetes deployment and monitoring (preferred).
  • Advanced ITIL certification or training, including ITIL v3 or v4 (preferred).
  • Experience integrating AI/ML into ITSM practices (preferred).

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Site Reliability Engineer

Counterpart Health 51-200 hospital & health care

Counterpart Health is hiring a Senior Site Reliability and Infrastructure Engineer to support and evolve the technology platform behind its primary care tool and maintain reliable infrastructure for domestic and international workloads.

AWS Azure CI/CD Containerd DNS Docker GCP Go gRPC Helm Kubernetes Linux Load Balancing Prometheus Python Shell Scripting TCP/IP
1 day, 20 hours ago

Senior Test Platform & Reliability Engineer - Star Trek Fleet Command

Scopely 1K-5K Internet Software & Services

Scopely is hiring a Senior Test Platform & Reliability Engineer in Ireland to build validation, reliability, and developer enablement platforms for Star Trek Fleet Command’s large-scale live-service backend systems.

AWS Bash CI/CD Docker GitLab Go Python Terraform
1 day, 20 hours ago

Senior Software Engineer - Databases, SRE | Canada | Remote

Grafana 1K-5K IT Services

Grafana Labs is hiring a Senior Software Engineer for its remote SRE team to improve reliability and operability of Grafana Cloud database services for high-SLA customers across AWS, GCP, and Azure.

AWS Azure GCP Go Helm Java Kubernetes Linux Microservices Python Terraform
2 days, 19 hours ago

Senior Site Reliability Engineer

Semios 51-250 Food Products

Semios Group is hiring a Senior Site Reliability Engineer to help scale, secure, and improve the reliability of its global agricultural technology platform.

AWS Azure Bash Buildkite CI/CD Datadog Docker Envoy GCP Git GitHub GitHub Actions GitLab Go Jenkins Kubernetes Linux NATS New Relic Prometheus Python Ruby Splunk Terraform
2 days, 21 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers