Senior/Principal Local LLM & Generative AI Platform Engineer

8 hours, 35 minutes ago
Full-time
Senior
DevOps and Infrastructure
Parallel Wireless

Parallel Wireless

Parallel Wireless leads the OpenRAN movement with innovative cloud-native architecture, enabling cost-effective cellular network deployment globally.

Wireless Telecommunication Services
251-1K
Founded 2012
$2M raised

Description

  • Own the architecture and technical roadmap for a secure, reliable local LLM platform.
  • Partner with engineering, product, support, IT, security, legal, and domain experts to prioritize use cases and define measurable requirements.
  • Build modular inference and model-gateway services with stable APIs, model routing, streaming, concurrency controls, and quotas.
  • Evaluate, select, and optimize open-weight language, code, embedding, reranking, and multimodal models.
  • Design and operate RAG and enterprise-search pipelines with metadata, embeddings, hybrid retrieval, reranking, citations, freshness, and deletion controls.
  • Enforce permissions across ingestion and retrieval, integrating identity, SSO, RBAC, secrets management, and audit logging.
  • Establish automated evaluation, regression testing, release gates, canary deployments, rollback procedures, and quality metrics.
  • Implement observability, production operations, CI/CD, registries, backups, disaster recovery, capacity planning, and incident response.
  • Design secure tool-calling and agent workflows with least privilege, sandboxing, validation, bounded execution, and human approval.
  • Integrate the platform into developer, source-control, CI, knowledge, ticketing, and internal application workflows through APIs and SDKs.

Requirements

  • BSc or MSc in Computer Science, Computer Engineering, Electrical Engineering, Data Science, or a related field, or equivalent practical experience.
  • Typically 7+ years of hands-on experience in software, ML platform, search, data, or infrastructure engineering, including recent LLM production experience.
  • Strong Python skills and experience building maintainable APIs, services, libraries, and data pipelines; Go, Java, or C/C++ experience is advantageous.
  • Strong understanding of transformer models and production inference, including tokenization, context management, batching, KV caching, parallelism, quantization, structured output, and tool calling.
  • Production experience building RAG or enterprise-search systems using embeddings, vector or lexical search, metadata filtering, reranking, source attribution, and retrieval evaluation.
  • Experience defining task-specific LLM evaluations using representative datasets, automated metrics, human feedback, error analysis, and regression thresholds.
  • Experience deploying containerized Linux services with Docker and Kubernetes or equivalent orchestration.
  • Experience with GPU-backed model serving, performance profiling, capacity planning, monitoring, reliability engineering, Git, testing, CI/CD, infrastructure as code, and incident response.
  • Strong knowledge of distributed systems, authentication, authorization, API security, secrets, encryption, auditability, and data lifecycle controls.
  • Preferred experience with private or air-gapped deployments, inference systems such as vLLM or TensorRT-LLM, GPU optimization, fine-tuning, vector databases, AI governance, LLM red-teaming, telecommunications/Open RAN, or relevant open-source projects.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Agentic AI Engineer

Weekday 11-50 Construction & Engineering

This full-time mid-senior software development role with one of our clients is remote within India and requires at least five years of experience.

8 hours, 50 minutes ago

Staff Engineer - Platform (Database & Data Infrastructure)

HighLevel 251-1K Internet Software & Services

HighLevel is seeking a Staff Engineer, Database & Data Infrastructure to define company-wide database strategy and architecture while improving scalability, reliability, developer experience, and infrastructure efficiency.

Cassandra ClickHouse CockroachDB Couchbase Databricks Druid DynamoDB Elasticsearch Firestore GCP Java Microservices MongoDB MySQL Node.js OpenSearch Oracle PostgreSQL Python Redis Snowflake SQL Server
8 hours, 50 minutes ago

ServiceNow Platform Engineer

Flexential 251-1K Internet Software & Services

Flexential is seeking a hands-on ServiceNow Engineer to design, implement, integrate, and support platform solutions in a data center and infrastructure-focused environment.

JavaScript
9 hours, 20 minutes ago

Senior/Lead Forward Deployed AI Engineer-Anthropic-US East

NewRocket 251-1K Internet Software & Services

NewRocket is seeking a Senior Forward Deployed AI Engineer for its AI Foundry team to work remotely with approximately 25% travel, delivering secure, production-ready AI workflows and integrations for enterprise customers while advancing NewRocket’s AI platforms and accelerators.

AWS Azure CI/CD Docker GCP Git JavaScript Microservices Python REST API Secrets Management SQL TypeScript
9 hours, 35 minutes ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers