Omilia

Omilia

Omilia is a global leader in Conversational AI, offering AI-based self-service solutions for enhanced customer care fulfillment and success.

IT Services
251-1K
Founded 2002
$20M raised

Description

  • Ensure platform reliability and availability across production and pre-production environments through proactive monitoring, alerting, and automation.
  • Serve as first response for incidents and contribute to problem management and root cause analysis.
  • Support development teams in building a reliability-focused culture within the development lifecycle.
  • Develop troubleshooting documentation and production support materials.
  • Collaborate with engineering teams to create optimized runbooks, operational documentation, and automation for operational tasks.
  • Work with development and cloud engineering teams to embed reliability and performance into the software delivery lifecycle.
  • Design, implement, and evolve observability solutions using metrics, logs, traces, and dashboards.
  • Use tools such as Prometheus, Grafana, and ELK to improve monitoring and visibility.
  • Participate in on-call rotations and continuously improve alert quality and response processes.
  • Champion continuous improvement in reliability and performance across teams.

Requirements

  • Bachelor's degree or MS in Engineering, or equivalent experience.
  • Experience operating at least one container orchestration cluster, such as Kubernetes or Docker Swarm.
  • Experience developing or maintaining software for production services at scale.
  • Experience with ELK.
  • Experience with AWS.
  • Experience with the Grafana/Prometheus stack.
  • Strong scripting skills in Bash, Python, or Go.
  • Excellent communication skills.
  • Ability to think creatively, anticipate challenges, and question existing technologies and procedures.
  • Comfort working in agile/lean methods and iterating collaboratively.
  • Strong team-player mindset and ability to work across product, experience design, engineering, and other functions.
  • Telephony knowledge, including SIP and VoIP, is a plus.
  • Experience in Linux administration, including RedHat, CentOS, or AL, is a plus.
  • Working knowledge of configuration management tools such as Terraform and Ansible is a plus.
  • Experience with TCP/IP and general networking concepts is a plus.
  • RDBMS knowledge, such as MySQL or Postgres, is a plus.
  • NoSQL knowledge, such as Redis, is a plus.

Benefits

  • Fixed compensation.
  • Long-term employment with vacation days.
  • Professional development support, including courses and training.
  • Opportunity to work on cutting-edge products with global impact in the service industry.
  • A collaborative, fun-to-work-with team.
  • Apple gear provided.
  • Equal opportunity employer with a diverse and inclusive workplace.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Cloud Performance Engineering - Site Reliability Engineer

Smile Digital Health 251-1K IT Services

Smile Digital Health is hiring a Cloud Site Reliability Engineer to ensure the reliability, scalability, and performance of production-grade healthcare data services across multiple cloud platforms.

Ansible Azure Chef Gatling Grafana HIPAA Java JMeter Kafka Kubernetes Linux Microservices OpenShift OpenTelemetry Prometheus Terraform
1 day, 4 hours ago

Senior Manager, Site Reliability Engineering

ClaritasRx 51-250 Pharmaceuticals

Claritas Rx is hiring a Senior Manager of Site Reliability Engineering to lead the reliability, cloud operations, and security posture of its AWS-hosted SaaS platform for healthcare and specialty drug access.

Apache Spark AWS AWS CDK CI/CD DynamoDB EC2 GitHub Actions Go HIPAA JIRA NestJS PostgreSQL Python React Secrets Management SQL Tableau Terraform TypeScript WAF
1 day, 4 hours ago

Senior Site Reliability | DevOps Engineer

Kaizen Gaming 1K-5K Hotels, Restaurants & Leisure

Kaizen Gaming is hiring a Senior Site Reliability Engineer to support the performance, stability, and scalability of large-scale applications and critical services across cloud and physical infrastructure.

Ansible Bash Chef CI/CD Docker Git GitLab CI Go Grafana Java Jenkins Kafka Kubernetes Logstash .NET PostgreSQL PowerShell Prometheus Python RabbitMQ Ruby SQL Server Terraform
2 days, 3 hours ago

Senior SRE / Production Reliability Engineer

Symphony Solutions 251-1K Internet Software & Services

Senior SRE role at an iGaming platform company responsible for building and maintaining reliable, observable, and scalable infrastructure across regulated-market products.

Argo CD CI/CD Couchbase DNS Docker Elasticsearch GCP GitOps Grafana Helm Kafka Kubernetes Linux Load Balancing Microservices OpsGenie PagerDuty PostgreSQL Prometheus React Scala Terraform TLS TypeScript
2 days, 4 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers