Back to work

Litmus

Red/Green CI for Prompt-Ware · Developer Tools

Tests for prompt-based skills: deterministic checks wherever possible, and a model judge only where it has been checked against known good and bad examples.

View source on GitHub

How it was built

Problem

Standard unit tests assume deterministic output, but a prompt-driven skill's behavior can vary run to run: a regular assertion either false-fails on harmless variation or gets loosened until it stops catching real regressions.

Decisions

  • Kept deterministic assertions for everything that is actually deterministic (inputs, outputs, file writes) rather than routing every check through a model.
  • Used a model judge only for the subjective parts, and refused to trust it until it correctly grades known pass and fail examples. A judge that says PASS to everything agrees with pass-only examples, so both kinds are required.
  • Designed it to run as a normal red/green suite inside existing CI, so a broken skill fails the build the same way a broken function does.

Outcome

Prompt-based skills get a red/green suite instead of "looked fine when I tried it once", and a judge can't turn a check green unless it has shown it can tell good output from bad.

Key features

  • Red/green test framework purpose-built for prompt-driven behavior, not just deterministic code
  • A judge check stays inconclusive until its rubric has both a passing and a failing example, and the judge agrees with them
  • Designed to run in CI alongside normal unit tests

Impact

Prompt-based skills get the same red/green CI discipline as regular code.

Tech stack

Python

Interested in something like this?

Get in touch