Senior AI Systems Quality Engineer

2 hours, 41 minutes ago
Full-time
Senior
Artificial Intelligence and Machine Learning
Abacus Insights

Abacus Insights

Abacus Insights simplifies healthcare data with intelligent solutions, unlocking data value and empowering health plans, consumers, and providers.

Insurance
51-250
Founded 2017
$82M raised

Description

  • Build production-grade validation frameworks, test harnesses, and evaluation pipelines across the AI lifecycle.
  • Design and evolve an AI testing platform integrated with Databricks and MLflow for repeatable testing, traceability, and auditability.
  • Create large-scale scenario-based test suites covering edge cases, long-tail scenarios, and failure modes.
  • Validate agent orchestration, including tool use, memory, decision logic, and non-deterministic behavior.
  • Define system contracts, guardrails, safe-degradation patterns, and measurable quality signals.
  • Integrate automated AI validation and quality gates into CI/CD pipelines for model, prompt, and code changes.
  • Develop reusable libraries and components for consistent AI quality practices.
  • Own aspects of AI release readiness and define go/no-go criteria using measurable thresholds.
  • Collaborate with AI, platform, security, domain, and delivery teams to establish production-relevant quality criteria.

Requirements

  • 7+ years of software engineering experience, primarily in backend or platform systems.
  • Production experience designing and implementing AI testing automation and custom validation frameworks.
  • Strong proficiency in Python and/or TypeScript.
  • Hands-on experience with LLM-based or agentic systems and non-deterministic behavior.
  • Experience with large-scale regression testing, long-tail evaluation, and broad test coverage.
  • Deep understanding of CI/CD integration and automated deployment quality gates.
  • Solid understanding of AWS cloud-native architectures.
  • Experience engineering for reliability, governance, safety, security, privacy, and operational risk in regulated or mission-critical environments.
  • Experience with AI evaluation methods such as drift detection, bias and fairness testing, hallucination assessment, and regression strategies.
  • Ability to define enforceable AI trust thresholds, including accuracy, hallucination limits, explainability, and PHI-safe behavior.
  • Experience working with domain experts to define correctness and production-relevant validation scenarios.
  • Preferred: Databricks Medallion architecture, MLflow, observability tools, LLM guardrails, performance and cost regression testing, AI/ML certifications, prompt and agent orchestration testing, or AI-generated test scenarios.

Benefits

  • Base salary plus eligibility for performance bonuses and equity grants.
  • Unlimited paid time off.
  • Work-from-anywhere flexibility.
  • Comprehensive health coverage with multiple plan options.
  • Equity for every employee.
  • Growth-focused professional development environment.
  • Home office setup allowance.
  • Monthly cell phone allowance.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Junior Automation Tester

Sparksoft 51-250 Internet Software & Services

Sparksoft is seeking a Test Automation and Quality Engineering professional to support government and federal healthcare applications by developing automated, performance, and regression testing solutions.

Agile AWS CI/CD DevSecOps Gatling Generative AI JavaScript Postman Python Scala Scrum Selenium SQL
1 hour, 41 minutes ago

GenAI Senior Integrated Designer

Brandtech+ 501-1000 Marketing services

Brandtech+ is seeking a remote GenAI Senior Integrated Designer to create and deliver digital, social, e-commerce, motion, and brand campaign assets using traditional design and GenAI workflows.

Digital Marketing Figma Generative AI Instagram API TikTok
2 hours, 41 minutes ago

QA Automation & Manual Engineer, Contract

DEPT® 1K-5K Media

DEPT® is seeking a Principal Engineer to establish agency-wide quality standards, lead scalable QA and performance testing practices, and provide technical oversight across client transformations.

Cypress JIRA JMeter Postman Python SQL TypeScript
2 hours, 41 minutes ago

AI Evaluation & Annotation Reviewer (L3 - Advanced Level) - Italian (US)

Volga Partners 51-250 Internet Software & Services

As an Italian AI Evaluation & Annotation Reviewer, you will support a global artificial intelligence project by evaluating and annotating AI-generated language content to improve response accuracy, relevance, and reliability.

Machine Learning
3 hours, 41 minutes ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers