Omilia

Omilia

Omilia is a global leader in Conversational AI, offering AI-based self-service solutions for enhanced customer care fulfillment and success.

IT Services
251-1K
Founded 2002
$20M raised

Description

  • Ensure platform reliability and availability across production and pre-production environments through proactive monitoring, alerting, and automation.
  • Serve as first response for incidents and contribute to problem management and root cause analysis.
  • Support the development team’s reliability efforts and help build a strong reliability culture within the development lifecycle.
  • Develop troubleshooting documentation for production support resources.
  • Collaborate with engineering teams to create runbooks, operational documentation, and automation for operational tasks.
  • Work with development and cloud engineering teams to embed reliability and performance into the software delivery lifecycle.
  • Design, implement, and evolve observability solutions using metrics, logs, traces, and dashboards.
  • Participate in on-call rotations and continuously improve alert quality and response processes.
  • Champion a culture of reliability, performance, and continuous improvement across teams.

Requirements

  • Bachelor’s degree or MS in Engineering, or equivalent experience.
  • Experience operating at least one container orchestration cluster, such as Kubernetes or Docker Swarm.
  • Experience developing or maintaining software for production services at scale.
  • Experience with ELK.
  • Experience with AWS.
  • Experience with the Grafana/Prometheus stack.
  • Strong scripting skills in Bash, Python, or Go.
  • Excellent communication skills.
  • Ability to think proactively, anticipate challenges, and challenge existing technologies, procedures, and thinking.
  • Versatility and willingness to iterate and learn in agile/lean environments.
  • Ability to work as a team player across product, design, engineering, and other functions.
  • Telephony knowledge, including SIP and VoIP, is a plus.
  • Experience in Linux administration, such as RedHat, CentOS, or AL, is a plus.
  • Working knowledge of configuration management tools such as Terraform and Ansible is a plus.
  • Experience with TCP/IP and general networking concepts is a plus.
  • RDBMS knowledge, such as MySQL or Postgres, is a plus.
  • NoSQL knowledge, such as Redis, is a plus.

Benefits

  • Fixed compensation.
  • Long-term employment with vacation days.
  • Professional development support, including courses and training.
  • Opportunity to work on cutting-edge technology products with global impact.
  • Proficient and fun-to-work-with colleagues.
  • Apple gear provided.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Site Reliability Engineer

PandaDoc 251-1K Internet Software & Services

PandaDoc is hiring a Site Reliability Engineer to help keep its document workflow platform reliable, resilient, and available for customers.

Agile AWS Django Grafana Java Kafka Kubernetes NATS PostgreSQL Python RabbitMQ Spring Boot
16 hours, 13 minutes ago

Senior Site Reliability Engineer

PandaDoc 251-1K Internet Software & Services

PandaDoc is hiring a Site Reliability Engineer to help keep its document workflow platform reliable through incident management, observability, production support, and resilience work across services.

Agile AWS Django Grafana Java Kafka Kubernetes NATS PostgreSQL Python RabbitMQ Spring Boot
16 hours, 13 minutes ago

Senior Site Reliability Engineer

PandaDoc 251-1K Internet Software & Services

PandaDoc is hiring a Site Reliability Engineer to keep its document workflow platform highly available and resilient while supporting production operations and reliability improvements.

Agile AWS Django Grafana Java Kafka Kubernetes NATS PostgreSQL Python RabbitMQ Spring Boot
1 day, 15 hours ago

Senior Site Reliability Engineer

PandaDoc 251-1K Internet Software & Services

PandaDoc is hiring a Site Reliability Engineer to help keep its production document workflow platform reliable, resilient, and low-downtime for customers.

AWS Django Grafana Java Kafka Kubernetes NATS PostgreSQL Python RabbitMQ Spring Boot
1 day, 15 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers