Omilia

Omilia

Omilia is a global leader in Conversational AI, offering AI-based self-service solutions for enhanced customer care fulfillment and success.

IT Services
251-1K
Founded 2002
$20M raised

Description

  • Ensure reliability and availability across production and pre-production environments through proactive monitoring, alerting, and automation.
  • Act as first response for incidents and contribute to problem management and root cause analysis.
  • Support development teams in improving service reliability and building a reliability-focused culture in the software lifecycle.
  • Develop troubleshooting documentation and operational runbooks for production support.
  • Collaborate with engineering and cloud teams to automate operational tasks and improve delivery processes.
  • Design, implement, and evolve observability solutions using metrics, logs, traces, and dashboards.
  • Use tools such as Prometheus, Grafana, and ELK to monitor platform health and performance.
  • Participate in on-call rotations and improve alert quality and incident response processes.
  • Champion continuous improvement in reliability, performance, and operational practices across teams.

Requirements

  • Bachelor’s degree or MS in Engineering, or equivalent experience.
  • Experience operating at least one container orchestration cluster such as Kubernetes or Docker Swarm.
  • Experience developing or maintaining software for production services at scale.
  • Experience with ELK.
  • Experience with AWS.
  • Experience with Grafana and Prometheus.
  • Strong scripting skills in Bash, Python, or Go.
  • Excellent communication skills and ability to work collaboratively across teams.
  • Agile/lean mindset with a willingness to iterate, learn, and challenge existing approaches.
  • Nice to have: telephony knowledge including SIP and VoIP.
  • Nice to have: Linux administration experience with RedHat, CentOS, or AL.
  • Nice to have: configuration management experience with Terraform or Ansible.
  • Nice to have: knowledge of TCP/IP and general networking concepts.
  • Nice to have: RDBMS experience with MySQL or Postgres.
  • Nice to have: NoSQL experience with Redis.

Benefits

  • Fixed compensation.
  • Long-term employment with vacation days.
  • Professional development support including courses and training.
  • Opportunity to work on cutting-edge technology products with global impact.
  • Collaborative and fun team environment.
  • Apple gear provided.
  • Equal opportunity employer with a diverse and inclusive workplace.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Site Reliability Engineer

PENN Entertainment 10K-50K Hotels, Restaurants & Leisure

PENN Entertainment is seeking a Senior Site Reliability Engineer to build and operate the cloud infrastructure supporting large-scale sports betting and media platforms across regulated jurisdictions.

Argo CD AWS Bash CI/CD Datadog GCP GitHub Actions GitOps Go Helm Kubernetes Linux PostgreSQL Python Shell Scripting Terraform
16 hours, 54 minutes ago

Staff Site Reliability Engineer, Ads

Reddit 1K-5K Internet Software & Services

Reddit is hiring a Staff Site Reliability Engineer to provide technical leadership for reliability, scalability, and operational excellence across its advertising infrastructure and revenue-critical systems.

Apache Spark ClickHouse GCP Go Kafka Kubernetes
17 hours, 24 minutes ago

Sr Lead Network Reliability Engineer

Coupa Software 1K-5K Internet Software & Services

Coupa is hiring a Sr. Lead Network Development Engineer to scale and operate its global SaaS platform’s cloud networking infrastructure through automation, reliability engineering, and technical leadership.

Ansible AWS Azure Chef DNS Fortinet Go Java Kubernetes Linux Python Ruby TCP/IP Terraform TLS
1 day, 17 hours ago

Senior Monitoring/Observability Architect

Makpar 51-250 Internet Software & Services

Makpar is seeking a Senior Monitoring/Observability Architect to lead enterprise monitoring strategy, architecture, and implementation guidance for a large federal government program.

Datadog Splunk
1 day, 17 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers