Back to work
Litmus
Red/Green CI for Prompt-Ware · Developer Tools
Tests for prompt-based skills: deterministic checks wherever possible, and a model judge only where it has been checked against known good and bad examples.
View source on GitHubHow it was built
Problem
Standard unit tests assume deterministic output, but a prompt-driven skill's behavior can vary run to run: a regular assertion either false-fails on harmless variation or gets loosened until it stops catching real regressions.
Decisions
- Kept deterministic assertions for everything that is actually deterministic (inputs, outputs, file writes) rather than routing every check through a model.
- Used a model judge only for the subjective parts, and refused to trust it until it correctly grades known pass and fail examples. A judge that says PASS to everything agrees with pass-only examples, so both kinds are required.
- Designed it to run as a normal red/green suite inside existing CI, so a broken skill fails the build the same way a broken function does.
Outcome
Prompt-based skills get a red/green suite instead of "looked fine when I tried it once", and a judge can't turn a check green unless it has shown it can tell good output from bad.
Key features
- Red/green test framework purpose-built for prompt-driven behavior, not just deterministic code
- A judge check stays inconclusive until its rubric has both a passing and a failing example, and the judge agrees with them
- Designed to run in CI alongside normal unit tests
Impact
Prompt-based skills get the same red/green CI discipline as regular code.
Tech stack
Python