Akkadian Labs

Akkadian Labs

Akkadian Labs is the top UC provisioning automation developer for Cisco Collaboration and Microsoft 365/Teams, offering a streamlined platform for MACD tasks with significant time and cost savings.

Internet Software & Services
51-250

Description

  • Support deployment and maintenance of scalable infrastructure in AWS and hybrid cloud environments.
  • Assist with infrastructure-as-code using Terraform, CloudFormation, or similar tools.
  • Maintain Linux-based environments and support containerization with Docker and orchestration with Kubernetes.
  • Design, deploy, and manage AI agent workloads, including compute provisioning and resource scaling for inference-heavy tasks.
  • Build and maintain model deployment pipelines, including versioning, testing, and rollback of AI models.
  • Monitor AI API consumption and infrastructure costs, and implement alerting and controls to prevent runaway usage.
  • Coordinate infrastructure-level security guardrails for AI systems, including access controls and data isolation.
  • Manage monitoring and observability using Prometheus, Grafana, and the ELK stack.
  • Troubleshoot system issues, support incident response, and perform root cause analysis.
  • Build, maintain, and optimize CI/CD pipelines and automate routine operational tasks such as builds, testing, deployments, and updates.
  • Collaborate with engineering, QA, product, and DevOps teams to support deployments, releases, and continuous improvement.
  • Maintain documentation for infrastructure, processes, and operational procedures.

Requirements

  • 5+ years of experience in DevOps, Site Reliability Engineering (SRE), or a related role.
  • Hands-on experience with AWS services such as EC2, ECS, S3, IAM, Lambda, and CloudWatch.
  • Working knowledge of Linux environments.
  • Familiarity with Docker and Kubernetes.
  • Basic to intermediate scripting ability in Python, Bash, or similar languages.
  • Experience building or maintaining CI/CD pipelines and related tools.
  • Exposure to monitoring and observability tools such as Prometheus, Grafana, and ELK.
  • Understanding of secure DevOps practices and basic compliance concepts.
  • Experience supporting AI or machine learning workloads and compute environments (preferred).
  • Exposure to AI model deployment pipelines and model versioning practices (preferred).
  • Experience with infrastructure-as-code tools such as Terraform or CloudFormation (preferred).
  • Familiarity with hybrid cloud or on-premises environments (preferred).
  • Exposure to security best practices in DevOps contexts, including AI-specific concerns such as data isolation and access controls (preferred).
  • Experience supporting production systems and participating in on-call rotations (preferred).

Benefits

  • Fully remote work environment.
  • Competitive benefits package.
  • Medical, dental, and vision insurance.
  • Company-paid life insurance and disability coverage.
  • 401(k) with a generous matching program.
  • Paid time off.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

PnP Pipeline On-call Support Engineer

PHIZENIX 11-50 information technology & services

This role at an internal engineering organization focuses on monitoring and triaging PnP (Power and Performance) pipelines to keep post-silicon work stable and moving efficiently.

CI/CD Jest
21 hours, 24 minutes ago

Senior Full Stack Engineer

Virtru 51-250 IT Services

Virtru is hiring a Senior Full Stack Engineer to build and operate its digital privacy products and core data protection platform for organizations that manage sensitive data.

AWS Docker GCP Go JavaScript Node.js PagerDuty React Selenium
2 days, 20 hours ago

Staff Forward Deployed Engineer

Tenstorrent 251-1K Internet Software & Services

Tenstorrent is hiring a remote Forward Deployed Engineer in North America to work directly with customers and internal teams on production AI inference deployments for its AI computers.

Grafana Helm Kubernetes OpenTelemetry Prometheus
2 days, 20 hours ago

Senior DevOps / MLOps Engineer (AI Agents, Claude)

Xebia 1K-5K Internet Software & Services

Xebia is hiring DevOps/MLOps Engineers for an international AI platform project focused on building the operational foundations for autonomous AI agent applications in the .NET ecosystem.

Azure CI/CD Docker Generative AI Linux LLM MLOps .NET Secrets Management
2 days, 21 hours ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers