Mozn

Mozn

MOZN is an enterprise AI company that has helped 100+ organizations make critical and informed decisions through specialized AI, in two key areas: Financial Crime Prevention and Enterprise Knowledge Intelligence

Internet Software & Services
51-250

Description

  • Participate in the on-call rotation, investigate and resolve incidents, perform root cause analysis, and document outcomes.
  • Debug application code and business logic, and ship fixes or pull requests when reliability issues belong in service repositories.
  • Design, build, and deploy LLM-based agents integrated with Kubernetes, cloud APIs, observability platforms, PagerDuty, and Slack.
  • Define agent tool interfaces and create secure wrappers for APIs, scripts, and read/write production actions.
  • Establish autonomy boundaries, approval gates, rollback paths, and human-in-the-loop controls for each agent.
  • Define agent correctness and safety criteria and build evaluation and backtesting suites using historical incidents.
  • Tune prompts, context, and tool schemas as agent capabilities and scope expand.
  • Collaborate with SRE and platform teams to identify repetitive, auditable workflows suitable for automation.
  • Measure agent impact through MTTD, MTTR, MTTX, error rates, and engineer-hours of toil removed.
  • Maintain security and compliance controls, including audit trails, least-privilege production access, and Saudi data residency requirements.

Requirements

  • 3+ years of experience building production software with LLMs, including agentic workflows, tool or function calling, multi-step planning, and RAG.
  • Hands-on experience delivering production work with Claude Code, OpenAI Codex, Kimi K2/K3, or a comparable agentic coding tool.
  • Strong Python or similar programming skills for agent tooling, API wrappers, and orchestration.
  • Hands-on SRE experience as an on-call responder, including incident response and root cause analysis.
  • Ability to read and debug application code and trace failures to the underlying logic before shipping fixes.
  • Experience with Kubernetes, a cloud provider such as AWS, GCP, OCI, or Azure, and observability tools including Prometheus, Grafana, Datadog, or ELK.
  • Understanding of autonomous-system guardrails, permissioning, approval gates, rollback procedures, and auditability.
  • Ability to build trust with technical stakeholders and increase agent autonomy responsibly.
  • Experience in Saudi Arabia or the MENA region, ideally consulting for public- or private-sector clients, is preferred.
  • Familiarity with Terraform, Ansible, Docker, virtual machines, on-premises environments, LLM agent evaluation, ML engineering, LLMOps, or platform engineering is preferred.

Benefits

  • Competitive compensation and top-tier health insurance.
  • High responsibility, autonomy, and trust in decision-making.
  • Collaborative workplace alongside AI specialists.
  • Inclusive culture that supports individual differences and professional growth.
  • Opportunity to work at a high-growth enterprise AI company in the Middle East.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Site Reliability Engineer

GoDaddy 5001-10000 Technology, Information and Internet

GoDaddy is hiring a Senior Site Reliability Engineer to support the Commerce ecosystem remotely by improving the reliability, scalability, security, and operation of business-critical production platforms.

Ansible AWS AWS CDK CI/CD CloudFormation Go Kubernetes Linux Pulumi Python SaltStack Terraform TypeScript
1 hour, 35 minutes ago

Senior Site Reliability Engineer

PENN Entertainment 10K-50K Hotels, Restaurants & Leisure

PENN Entertainment is seeking a Senior Site Reliability Engineer to build and operate the cloud infrastructure supporting large-scale sports betting and media platforms across regulated jurisdictions.

Argo CD AWS Bash CI/CD Datadog GCP GitHub Actions GitOps Go Helm Kubernetes Linux PostgreSQL Python Shell Scripting Terraform
1 day, 1 hour ago

Staff Site Reliability Engineer, Ads

Reddit 1K-5K Internet Software & Services

Reddit is hiring a Staff Site Reliability Engineer to provide technical leadership for reliability, scalability, and operational excellence across its advertising infrastructure and revenue-critical systems.

Apache Spark ClickHouse GCP Go Kafka Kubernetes
1 day, 1 hour ago

Sr Lead Network Reliability Engineer

Coupa Software 1K-5K Internet Software & Services

Coupa is hiring a Sr. Lead Network Development Engineer to scale and operate its global SaaS platform’s cloud networking infrastructure through automation, reliability engineering, and technical leadership.

Ansible AWS Azure Chef DNS Fortinet Go Java Kubernetes Linux Python Ruby TCP/IP Terraform TLS
2 days, 1 hour ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers