Mozn

Mozn

MOZN is an enterprise AI company that has helped 100+ organizations make critical and informed decisions through specialized AI, in two key areas: Financial Crime Prevention and Enterprise Knowledge Intelligence

Internet Software & Services
51-250

Description

  • Participate in the on-call rotation, investigate and resolve incidents, perform root cause analysis, and document outcomes.
  • Debug application code and business logic, and ship fixes or pull requests when reliability issues belong in service repositories.
  • Design, build, and deploy LLM-based agents integrated with Kubernetes, cloud APIs, observability platforms, PagerDuty, and Slack.
  • Define agent tool interfaces and create secure wrappers for APIs, scripts, and read/write production actions.
  • Establish autonomy boundaries, approval gates, rollback paths, and human-in-the-loop controls for each agent.
  • Define agent correctness and safety criteria and build evaluation and backtesting suites using historical incidents.
  • Tune prompts, context, and tool schemas as agent capabilities and scope expand.
  • Collaborate with SRE and platform teams to identify repetitive, auditable workflows suitable for automation.
  • Measure agent impact through MTTD, MTTR, MTTX, error rates, and engineer-hours of toil removed.
  • Maintain security and compliance controls, including audit trails, least-privilege production access, and Saudi data residency requirements.

Requirements

  • 3+ years of experience building production software with LLMs, including agentic workflows, tool or function calling, multi-step planning, and RAG.
  • Hands-on experience delivering production work with Claude Code, OpenAI Codex, Kimi K2/K3, or a comparable agentic coding tool.
  • Strong Python or similar programming skills for agent tooling, API wrappers, and orchestration.
  • Hands-on SRE experience as an on-call responder, including incident response and root cause analysis.
  • Ability to read and debug application code and trace failures to the underlying logic before shipping fixes.
  • Experience with Kubernetes, a cloud provider such as AWS, GCP, OCI, or Azure, and observability tools including Prometheus, Grafana, Datadog, or ELK.
  • Understanding of autonomous-system guardrails, permissioning, approval gates, rollback procedures, and auditability.
  • Ability to build trust with technical stakeholders and increase agent autonomy responsibly.
  • Experience in Saudi Arabia or the MENA region, ideally consulting for public- or private-sector clients, is preferred.
  • Familiarity with Terraform, Ansible, Docker, virtual machines, on-premises environments, LLM agent evaluation, ML engineering, LLMOps, or platform engineering is preferred.

Benefits

  • Competitive compensation and top-tier health insurance.
  • High responsibility, autonomy, and trust in decision-making.
  • Collaborative workplace alongside AI specialists.
  • Inclusive culture that supports individual differences and professional growth.
  • Opportunity to work at a high-growth enterprise AI company in the Middle East.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Site Reliability Engineer - AWS and Azure

Jalasoft 1K-5K Internet Software & Services

Jalasoft is hiring a Site Reliability Engineer to build, maintain, and improve reliable, scalable, and secure cloud infrastructure across Windows and Linux environments.

Ansible AWS Azure Bash CI/CD Docker GitHub GitHub Actions Grafana Kubernetes PowerShell Prometheus Python TeamCity Terraform
15 hours, 44 minutes ago

[Job-31445] Sênior Software Engineer | SRE & Software Architecture, Brazil

CI&T 5K-10K Internet Software & Services

A CI&T busca uma pessoa especialista em observabilidade para atuar como facilitadora técnica junto aos times de desenvolvimento, apoiando a confiabilidade, a performance e a evolução das aplicações.

Agile Angular Azure CI/CD Datadog Docker Git GitFlow GitHub Actions Grafana Java Kanban Kubernetes OpenShift OpenTelemetry Prometheus Scrum WAF
1 day, 15 hours ago

Senior Site Reliability Engineer

Sports Academy Education Services

Texas Sports Academy is seeking a part-time Senior Site Reliability Engineer consultant to audit, improve, and scale the infrastructure supporting its AI-first K–12 school.

AWS CI/CD Datadog Grafana Prometheus
1 day, 15 hours ago

Site Reliability Engineer (SRE)

Rocket.net 11-50 IT Services

Rocket.net is seeking a Site Reliability Engineer to maintain the reliability and performance of its hosting platform while resolving complex infrastructure issues and providing advanced support to customers.

Apache Bash CDN Cloudflare Datadog DNS Linux MariaDB MySQL Nginx Redis SSH WAF WordPress
3 days, 15 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers