AI Prompt Regression Kit

Catch prompt regressions before your users do

A prompt edit, a model bump, a 'small' system-prompt tweak — and behavior silently changes for everyone. Eval platforms charge $29–39 per seat, every month, to watch it happen in a dashboard. This kit is a gate instead: a golden set of prompt cases, a Python runner (stdlib only), LLM-as-judge rubrics and a GitHub Actions check that turns the offending PR red. Runs on your own API key, on any OpenAI-compatible endpoint. One payment of ฿990 — less than one month of the cheapest paid plan — then nothing.

Get the kit — ฿990
One-time payment · ≈ US$28 · free updates for v1.x · instant download

What's inside the kit

One zip, 14 files — no install, no account, no service to sign up for. Everything runs on your machine, with Python 3 only.

The product

eval-runner.py

Runs your cases.csv against any OpenAI-compatible endpoint, scores deterministic assertions plus LLM-as-judge rubrics, and writes JSON + Markdown reports with run-to-run diffs. Python 3 standard library only — nothing to pip install.

Start here

Golden-set template

A starter set of prompt cases in plain CSV — input, expected behavior, assertions — so your first 20 regression cases are an afternoon of work, not a project.

Score what strings can't

5 judge rubrics

Five LLM-as-judge rubrics with 1–5 anchors (tone, faithfulness, language consistency and more), plus docs on writing your own and keeping judge variance under control.

The gate

GitHub Actions workflow

A ready workflow that runs your eval on every PR and fails it when the score drops — plus an offline smoke job. The PR that makes the bot rude goes red.

The differentiator

Multilingual eval guide

The EN/JA/PT-BR/DE mirror-case method: catch the silent failure class where Japanese users get English answers — the regression every major tool treats as an afterthought.

60 seconds to proof

Working sample project

A 5-case sample (EN/JA/DE) with stored responses that runs offline — no API key, no network — and shows GATE PASSED and a diff before you spend a cent on API calls.

How it works

Three steps from 'poked it in the playground' to a real regression gate.

Step 1

Write your cases

Add the prompts and inputs your users actually hit to the golden-set CSV, with assertions for what must be true and rubrics for what only a judge can grade.

Step 2

Run the eval

eval-runner.py calls your endpoint — cloud or localhost — and writes a JSON + Markdown report with pass/fail per case and a diff against the last run.

Step 3

Gate your PRs

Drop github-actions-gate.yml into .github/workflows. From then on, every pull request that degrades eval quality fails CI — before it merges.

Runs on your machine, your key, your rules

No data leaves your machine — except to the endpoint you configure. The runner calls your model with your API key, straight from your laptop or CI runner. No account, no server of ours in the path, no telemetry, no golden set uploaded anywhere.

Local models welcome. Any OpenAI-compatible base URL works — including Ollama, LM Studio and vLLM on localhost. Same gate for llama as for gpt.

Honest scope: this is a regression gate over a curated golden set, not a tracing platform — no production capture, no dashboards. CI runs cost API money per PR (a 30-case set is pennies, not free).

One price. It's yours.

No subscription, no seats, no usage meter.

฿990 one-time payment

≈ US$28

  • Free updates for v1.x
  • Instant download after checkout
  • Single developer/team commercial license
  • Any OpenAI-compatible endpoint — no lock-in

Secure checkout via Stripe · card payments · taxes calculated at checkout

The cheapest paid tier of an eval platform is $29/month (Langfuse Core, verified Sep 2026); LangSmith Plus is $39/seat/month. One kit ≈ one of those months — paid once.

Frequently asked questions

What does the license cover?

A single commercial license for one developer or one team: the buyer can use the kit on unlimited projects, and the whole team shares the same zip. Redistributing or reselling the kit files themselves is not allowed. Every golden set, report and gate you build with it is entirely yours.

Do I get updates?

Yes — every v1.x update (new rubrics, runner improvements, guide revisions) is free. Re-download any time from your download page with your email and password. If a v2.0 ever ships, existing buyers get a clear upgrade path — never a surprise paywall.

What is the refund policy?

If the kit doesn't fit how you work, email zicula06@gmail.com within 14 days of purchase and we'll refund you in full — the zip is yours to keep or delete. One refund per purchase; we only ask for optional feedback.

Which LLM providers does it work with?

Any endpoint that speaks the OpenAI chat-completions format: OpenAI, Azure OpenAI, OpenRouter, Groq, Together, Mistral — and local servers like Ollama, LM Studio or vLLM. Point the runner's --base-url at your endpoint and run; no code changes.

Do I need a paid API key to try it?

No. The kit includes an offline demo mode and a working sample project that runs with no key and no network — you'll see GATE PASSED and a run-to-run diff in about a minute. Live eval calls use your own key and cost pennies for a typical 30-case set.

Ship your next prompt change with a green gate

฿990 one-time · ≈ US$28 · free updates v1.x · instant download

Get the AI Prompt Regression Kit — ฿990