Tekion

Tekion

Tekion is a leading provider of cloud-native automotive platforms that unify DMS, CRM, Digital Retail, Analytics, and more. Their AI-powered software enables personalized selling, upsell, and cross-sell opportunities, driving revenue and profitability....

IT Services
1K-5K
Founded 2016
$435M raised

Description

  • Build automated testing suites to detect hallucinations, bias, toxicity, and prompt injection vulnerabilities in LLM-powered products.
  • Implement automated evaluations for RAG systems, including context relevance, groundedness, and answer faithfulness.
  • Design test beds for validating multi-agent workflows, including tool-calling accuracy, reasoning, memory, and decision loops.
  • Build and run automated conversation simulations to stress-test agent behavior across intents, edge cases, and multi-turn flows.
  • Create prompt regression frameworks to measure the impact of prompt and parameter changes on output consistency.
  • Statistically validate AI data outputs to identify silent quality failures before production.
  • Audit data ingestion, transformation, and feature store pipelines for schema drift and data corruption.
  • Validate vector database indexing, embedding similarity accuracy, and retrieval latency.
  • Maintain automated suites for ML metrics and deep learning loss curves across model versions.
  • Embed evaluation and data QA checks into MLOps and CI/CD pipelines so quality failures block releases automatically.

Requirements

  • 8+ years of experience in an SDET role.
  • Expert-level Python experience for test automation, evaluation pipelines, and data analysis.
  • Strong SQL experience for data output validation, ground truth querying, and pipeline data quality checks.
  • Experience with LLM evaluation frameworks such as RAGAS, TruLens, DeepEval, or Promptflow.
  • Experience with agent workflow testing and LLM debugging tools such as LangChain, LangSmith, or LlamaIndex.
  • Experience with OpenAI, Anthropic, or Hugging Face APIs.
  • Experience with vector databases and retrieval quality testing.
  • Experience with Pandas and NumPy for statistical analysis, data profiling, and schema validation.
  • Experience with Postman, REST Assured, or Requests for API contract validation and integration testing.
  • Experience with MLflow and Docker, GitHub Actions, or Jenkins for CI/CD automation.
  • Good-to-have experience with AWS Bedrock, Azure OpenAI, or GCP Vertex AI.
  • Good-to-have experience with Kubeflow, Weights & Biases, or Feast.
  • Good-to-have familiarity with Scikit-learn, TensorFlow, or PyTorch.
  • Good-to-have experience with Kubernetes and Terraform.
  • Good-to-have experience with Playwright or Cypress.
  • Good-to-have experience with Locust or JMeter.
  • Good-to-have knowledge of statistical hypothesis testing, including t-tests, confidence intervals, and significance testing.
  • Good-to-have experience generating synthetic data with LLMs.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Shape the Future of AI — Ilocano Talent Hub

Welo Global Professional Services

Welo Data, part of Welocalize, is recruiting a global remote contributor network for short-term AI data projects that support annotation, evaluation, and prompt creation for safer, more accurate AI.

18 minutes ago

Project Epsilon - French (France) Data Trainer

Welo Global Professional Services

Welo Data is hiring remote freelance Trainers for a long-term AI data annotation project to review image-based questions and provide accurate answers from visual evidence.

LLM
18 minutes ago

Project Epsilon - Dutch Data Trainer

Welo Global Professional Services

Welo Data is hiring remote freelance Trainers to support an AI data annotation project by reviewing image-based questions and providing accurate golden answers.

18 minutes ago

Project Epsilon - Korean Data Trainer

Welo Global Professional Services

Welo Data, part of Welocalize, is hiring remote freelance Trainers to support an AI data annotation project by reviewing image-based questions and providing accurate golden answers.

18 minutes ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers