Omilia

Omilia

Omilia is a global leader in Conversational AI, offering AI-based self-service solutions for enhanced customer care fulfillment and success.

IT Services
251-1K
Founded 2002
$20M raised

Description

  • Ensure reliability and availability across production and pre-production environments through proactive monitoring, alerting, and automation.
  • Act as first response for incidents and contribute to problem management and root cause analysis.
  • Support development teams in improving service reliability and building a reliability-focused culture in the software lifecycle.
  • Develop troubleshooting documentation and operational runbooks for production support.
  • Collaborate with engineering and cloud teams to automate operational tasks and improve delivery processes.
  • Design, implement, and evolve observability solutions using metrics, logs, traces, and dashboards.
  • Use tools such as Prometheus, Grafana, and ELK to monitor platform health and performance.
  • Participate in on-call rotations and improve alert quality and incident response processes.
  • Champion continuous improvement in reliability, performance, and operational practices across teams.

Requirements

  • Bachelor’s degree or MS in Engineering, or equivalent experience.
  • Experience operating at least one container orchestration cluster such as Kubernetes or Docker Swarm.
  • Experience developing or maintaining software for production services at scale.
  • Experience with ELK.
  • Experience with AWS.
  • Experience with Grafana and Prometheus.
  • Strong scripting skills in Bash, Python, or Go.
  • Excellent communication skills and ability to work collaboratively across teams.
  • Agile/lean mindset with a willingness to iterate, learn, and challenge existing approaches.
  • Nice to have: telephony knowledge including SIP and VoIP.
  • Nice to have: Linux administration experience with RedHat, CentOS, or AL.
  • Nice to have: configuration management experience with Terraform or Ansible.
  • Nice to have: knowledge of TCP/IP and general networking concepts.
  • Nice to have: RDBMS experience with MySQL or Postgres.
  • Nice to have: NoSQL experience with Redis.

Benefits

  • Fixed compensation.
  • Long-term employment with vacation days.
  • Professional development support including courses and training.
  • Opportunity to work on cutting-edge technology products with global impact.
  • Collaborative and fun team environment.
  • Apple gear provided.
  • Equal opportunity employer with a diverse and inclusive workplace.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Site Reliability Engineer

Counterpart Health 51-200 hospital & health care

Counterpart Health is hiring a Senior Site Reliability and Infrastructure Engineer to support and evolve the technology platform behind its primary care tool and maintain reliable infrastructure for domestic and international workloads.

AWS Azure CI/CD Containerd DNS Docker GCP Go gRPC Helm Kubernetes Linux Load Balancing Prometheus Python Shell Scripting TCP/IP
16 hours, 7 minutes ago

Senior Test Platform & Reliability Engineer - Star Trek Fleet Command

Scopely 1K-5K Internet Software & Services

Scopely is hiring a Senior Test Platform & Reliability Engineer in Ireland to build validation, reliability, and developer enablement platforms for Star Trek Fleet Command’s large-scale live-service backend systems.

AWS Bash CI/CD Docker GitLab Go Python Terraform
16 hours, 22 minutes ago

Senior Software Engineer - Databases, SRE | Canada | Remote

Grafana 1K-5K IT Services

Grafana Labs is hiring a Senior Software Engineer for its remote SRE team to improve reliability and operability of Grafana Cloud database services for high-SLA customers across AWS, GCP, and Azure.

AWS Azure GCP Go Helm Java Kubernetes Linux Microservices Python Terraform
1 day, 15 hours ago

Senior Site Reliability Engineer

Semios 51-250 Food Products

Semios Group is hiring a Senior Site Reliability Engineer to help scale, secure, and improve the reliability of its global agricultural technology platform.

AWS Azure Bash Buildkite CI/CD Datadog Docker Envoy GCP Git GitHub GitHub Actions GitLab Go Jenkins Kubernetes Linux NATS New Relic Prometheus Python Ruby Splunk Terraform
1 day, 16 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers