Tekion

Tekion

Tekion is a leading provider of cloud-native automotive platforms that unify DMS, CRM, Digital Retail, Analytics, and more. Their AI-powered software enables personalized selling, upsell, and cross-sell opportunities, driving revenue and profitability....

IT Services
1K-5K
Founded 2016
$435M raised

Description

  • Build automated testing suites to detect hallucinations, bias, toxicity, and prompt injection vulnerabilities in LLM-powered products.
  • Implement automated evaluations for RAG systems to measure context relevance, groundedness, and answer faithfulness.
  • Design test beds for multi-agent workflows, including tool-calling accuracy, multi-step reasoning, memory, and autonomous decision loops.
  • Build and run automated conversation simulations to stress-test agent behavior across intents, edge cases, and multi-turn dialogs.
  • Create prompt regression frameworks to assess how changes in prompts, temperature, and sampling parameters affect output consistency.
  • Statistically validate AI data outputs using distributions, precision/recall, and error pattern analysis to catch data quality failures before production.
  • Audit data ingestion, transformation, and feature store pipelines for schema drift and data corruption.
  • Validate vector database indexing, embedding similarity accuracy, and retrieval latency.
  • Maintain automated suites for ML metrics and deep learning loss curves across model versions.
  • Embed AI evaluation and data QA checks into MLOps and CI/CD pipelines so quality failures block releases automatically.

Requirements

  • 5+ years of experience in an SDET role.
  • Expert-level Python experience for test automation, evaluation pipelines, and data analysis.
  • Strong SQL experience for data output validation, ground-truth querying, and pipeline quality checks.
  • Experience with RAGAS, TruLens, DeepEval, or Promptflow for LLM evaluation.
  • Experience with LangChain, LangSmith, or LlamaIndex for agent workflow testing and prompt tracing.
  • Experience with OpenAI, Anthropic, or Hugging Face APIs for direct LLM endpoint testing.
  • Experience with vector databases for retrieval quality testing and latency benchmarking.
  • Experience with Pandas and NumPy for statistical analysis, profiling, schema validation, and pipeline integrity checks.
  • Experience with Pytest and API automation tools such as Postman, REST Assured, or Requests.
  • Familiarity with MLflow, Docker, GitHub Actions, Jenkins, Grafana, Kibana, and OpenTelemetry.
  • Preferred experience with AWS Bedrock, Azure OpenAI, or GCP Vertex AI.
  • Preferred experience with Kubeflow, Weights & Biases, or Feast feature stores.
  • Preferred familiarity with Scikit-learn, TensorFlow, or PyTorch.
  • Preferred experience with Kubernetes, Terraform, Playwright, Cypress, Locust, or JMeter.
  • Preferred knowledge of statistical hypothesis testing, including t-tests, confidence intervals, and significance testing.
  • Preferred experience generating synthetic data using LLMs.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Shape the Future of AI — Ilocano Talent Hub

Welo Global Professional Services

Welo Data, part of Welocalize, is recruiting a global remote contributor network for short-term AI data projects that support annotation, evaluation, and prompt creation for safer, more accurate AI.

18 minutes ago

Project Epsilon - French (France) Data Trainer

Welo Global Professional Services

Welo Data is hiring remote freelance Trainers for a long-term AI data annotation project to review image-based questions and provide accurate answers from visual evidence.

LLM
18 minutes ago

Project Epsilon - Dutch Data Trainer

Welo Global Professional Services

Welo Data is hiring remote freelance Trainers to support an AI data annotation project by reviewing image-based questions and providing accurate golden answers.

18 minutes ago

Project Epsilon - Korean Data Trainer

Welo Global Professional Services

Welo Data, part of Welocalize, is hiring remote freelance Trainers to support an AI data annotation project by reviewing image-based questions and providing accurate golden answers.

18 minutes ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers