563
Tool Stack

AI Testing Stack

Test deterministic code and probabilistic AI output in one pipeline.

1 min readupdated 2026-08-01

/ quick answer

Keep quality measurable when part of the system is non-deterministic. Test deterministic code and probabilistic AI output in one pipeline.

Test deterministic code and probabilistic AI output in one pipeline. Keep quality measurable when part of the system is non-deterministic. The stack combines Vitest or Pytest (unit and integration), Playwright (end-to-end), Langfuse or Braintrust (eval runs and scoring), GitHub Actions (gates on every commit), A stored regression set of real production failures. This tool stack node is part of the Onexial knowledge graph and links to related concepts, workflows and tools below.
Purpose
Keep quality measurable when part of the system is non-deterministic.
Tools Included
  • Vitest or Pytest (unit and integration)
  • Playwright (end-to-end)
  • Langfuse or Braintrust (eval runs and scoring)
  • GitHub Actions (gates on every commit)
  • A stored regression set of real production failures
Workflow Supported
Alternatives
  • Promptfoo
  • DeepEval
  • Custom eval harness
/ frequently asked

What is the AI Testing Stack stack for?

Keep quality measurable when part of the system is non-deterministic.

Which tools are in this stack?

Vitest or Pytest (unit and integration), Playwright (end-to-end), Langfuse or Braintrust (eval runs and scoring), GitHub Actions (gates on every commit), A stored regression set of real production failures.

Are there alternatives to this stack?

Yes — Promptfoo, DeepEval, Custom eval harness.

/ continue exploring

Related concepts

The vocabulary this page depends on.

  • AI Software Engineering

    AI software engineering is the practice of building software where agents write most of the code and humans own architecture, review and verification.

  • AI Testing

    AI testing covers two things: using AI to generate and maintain tests, and testing AI systems whose output is non-deterministic.

  • AI Evaluation

    AI evaluation is the measurement layer of an AI system: a fixed set of cases, a scoring method and a tracked pass rate you can regress against.

all dictionary

Related workflows

Turn this into a repeatable process.

all workflows

Related tool stacks

The tools that run it in production.

all tool stacks

Related prompts

Reusable prompts for this job.

all prompts

Related use cases

How people apply it, and what came out.

all use cases

Comparisons & alternatives

Pick between the options.

all comparisons