Mozn

Mozn

MOZN is an enterprise AI company that has helped 100+ organizations make critical and informed decisions through specialized AI, in two key areas: Financial Crime Prevention and Enterprise Knowledge Intelligence

Internet Software & Services
51-250

Description

  • Participate in the on-call rotation, investigate and resolve incidents, perform root cause analysis, and document outcomes.
  • Debug application code and business logic, and ship fixes or pull requests when reliability issues belong in service repositories.
  • Design, build, and deploy LLM-based agents integrated with Kubernetes, cloud APIs, observability platforms, PagerDuty, and Slack.
  • Define agent tool interfaces and create secure wrappers for APIs, scripts, and read/write production actions.
  • Establish autonomy boundaries, approval gates, rollback paths, and human-in-the-loop controls for each agent.
  • Define agent correctness and safety criteria and build evaluation and backtesting suites using historical incidents.
  • Tune prompts, context, and tool schemas as agent capabilities and scope expand.
  • Collaborate with SRE and platform teams to identify repetitive, auditable workflows suitable for automation.
  • Measure agent impact through MTTD, MTTR, MTTX, error rates, and engineer-hours of toil removed.
  • Maintain security and compliance controls, including audit trails, least-privilege production access, and Saudi data residency requirements.

Requirements

  • 3+ years of experience building production software with LLMs, including agentic workflows, tool or function calling, multi-step planning, and RAG.
  • Hands-on experience delivering production work with Claude Code, OpenAI Codex, Kimi K2/K3, or a comparable agentic coding tool.
  • Strong Python or similar programming skills for agent tooling, API wrappers, and orchestration.
  • Hands-on SRE experience as an on-call responder, including incident response and root cause analysis.
  • Ability to read and debug application code and trace failures to the underlying logic before shipping fixes.
  • Experience with Kubernetes, a cloud provider such as AWS, GCP, OCI, or Azure, and observability tools including Prometheus, Grafana, Datadog, or ELK.
  • Understanding of autonomous-system guardrails, permissioning, approval gates, rollback procedures, and auditability.
  • Ability to build trust with technical stakeholders and increase agent autonomy responsibly.
  • Experience in Saudi Arabia or the MENA region, ideally consulting for public- or private-sector clients, is preferred.
  • Familiarity with Terraform, Ansible, Docker, virtual machines, on-premises environments, LLM agent evaluation, ML engineering, LLMOps, or platform engineering is preferred.

Benefits

  • Competitive compensation and top-tier health insurance.
  • High responsibility, autonomy, and trust in decision-making.
  • Collaborative workplace alongside AI specialists.
  • Inclusive culture that supports individual differences and professional growth.
  • Opportunity to work at a high-growth enterprise AI company in the Middle East.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Cloud Site Reliability Engineer

Cadwell 51-250 Health Care Providers & Services

Cadwell is seeking a Cloud Site Reliability Engineer to operate and improve AWS infrastructure supporting healthcare customers and ensure reliable, secure, and compliant hosted neurodiagnostic software environments.

AWS Bash CI/CD Encryption HIPAA JavaScript JSON Python Terraform TypeScript YAML
10 hours, 30 minutes ago

Site Reliability Engineer

GiveCampus 51-250 Internet Software & Services

GiveCampus is seeking a hands-on Site Reliability Engineer to strengthen the reliability, performance, observability, and operational maturity of its AWS-based fundraising platform in a remote-first U.S. role.

AWS CI/CD CircleCI Datadog GitHub Actions Kubernetes Linux New Relic OpenSearch PostgreSQL Redis Ruby Ruby on Rails Terraform
11 hours ago

DevOps / SRE / DevSecOps Engineer (AWS) - Latin America, Remote

Bluelight Consulting 11-50 Internet Software & Services

Bluelight is hiring a DevOps/SRE/DevSecOps professional to operate secure, reliable AWS-based systems and delivery infrastructure for client software projects in a remote consultancy environment.

AWS CI/CD DevSecOps Docker K6 OpenTelemetry PostgreSQL Secrets Management Terraform
1 day, 10 hours ago

DevOps / SRE / DevSecOps Engineer (AWS) - Latin America, Remote

Bluelight Consulting 11-50 Internet Software & Services

Bluelight is hiring a DevOps/SRE/DevSecOps professional to support complex client systems by building secure, reliable, and observable AWS infrastructure and delivery operations.

AWS CI/CD DevSecOps Docker K6 OpenTelemetry PostgreSQL Secrets Management Terraform
1 day, 10 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers