LLM Evaluation Pipelines in CI: The Engineer Setup

LLM evaluation pipelines in CI turn model and prompt quality from a manual review into a build gate: a versioned test set, a runner that executes it, and a threshold that fails the pipeline when scores drop. For most engineering teams the practical stack is two open-source tools. promptfoo evaluates prompts, agents and retrieval flows as ordinary CI tests, while EleutherAI’s lm-evaluation-harness scores hosted models against standardized academic benchmarks. promptfoo ships a single-command eval runner that emits JSON, HTML and JUnit XML output, so one results file can feed both your quality gate and your CI system’s native test viewer.

Why prompts need CI gates

A prompt file is code with no compiler. Nothing stops a one-line edit from breaking JSON output contracts, refusal behaviour or answer accuracy until a user notices in production. Provider-side model updates create the same risk without any commit in your repository, which is why evaluation has to run continuously rather than once per release. Retrieval changes amplify the blast radius: if you run semantic search, swapping an embedding model alters ranked results for every query at once, and the decision criteria for that swap belong next to the tests that catch it — the trade-offs we walked through for choosing embedding models for semantic search apply here too. A CI gate converts those silent regressions into a red build with a diff attached.

Choosing the eval harness

Match the tool to what actually changed in the commit. promptfoo is built for product surfaces: prompts, RAG pipelines, agents and tool-calling flows, with assertions ranging from exact-match and contains checks to model-graded rubrics and deterministic JSON schema validation. lm-evaluation-harness targets the model itself: it is the tool you reach for when you fine-tune, quantize or swap a self-hosted checkpoint and need comparable scores across runs.

ToolScopeOutput formatsBest fit
promptfooPrompts, agents, RAG, red-team scansJSON, HTML, JUnit XMLPull-request gates on product behaviour
lm-evaluation-harnessModel-level academic benchmarksJSON, CSV, Markdown tablesFine-tune, quantization and checkpoint comparisons
Custom harness scriptsBusiness KPIs, domain metricsWhatever your CI ingestsMetric definitions promptfoo does not expose

On the benchmark side, the harness implements more than 60 standard academic benchmarks and is the evaluation backend for Hugging Face’s Open LLM Leaderboard, which is what makes its scores comparable across teams and papers. It also stays practical at the infrastructure level: the harness supports fast and memory-efficient inference with vLLM alongside commercial API providers, so a CI runner can evaluate a self-hosted checkpoint at full throughput instead of trickling requests through a naive client. If you already serve models with vLLM, the same capacity levers covered in KV cache and batching for LLM serving throughput shorten nightly benchmark runs by a wide margin.

Designing the test suite

The suite is the asset; the runner is replaceable. Build it like this:

  • Golden set first: 50–200 real production cases with expected outcomes, stored in the repo and reviewed like code.
  • Assertion per failure mode: schema checks for structure, contains checks for citations, rubric grades for tone, refusal tests for safety — never one vague score for everything.
  • Per-test thresholds: exact-match for extraction tasks, similarity floors for generation, so one noisy test cannot sink the gate.
  • Pinned inputs: freeze model version, temperature and retrieval index snapshot per run, or comparisons across commits are meaningless.
  • Cost ceiling: a token budget per pipeline run that fails loudly when a prompt edit triples consumption.

Treat the suite as a regression net, not a leaderboard. Every production incident should end with a new test case added, which is how the suite compounds in value while the runner stays disposable.

The CI workflow, step by step

  1. Commit prompts, retrieval configs and the golden test set under prompts/ and tests/, with promptfooconfig.yaml at the root.
  2. Trigger evaluation only on pull requests that touch those paths, plus a scheduled nightly run for provider-side model drift.
  3. Run npx promptfoo@latest eval -o results.json -o results.junit.xml, publishing the JUnit file so failures render inside your CI system’s test tab.
  4. Gate the merge: fail on any error, or compute the pass rate from the JSON stats and block below your threshold (95% is a common floor).
  5. Cache evaluation results keyed on the config and prompt files, so unchanged prompts cost zero tokens on re-runs.

Check the runtime before the first run: promptfoo needs Node.js 22.22.0 or newer in the CI environment, and lists Node.js 24 LTS as the recommended runtime, so pin the Node version in the workflow instead of relying on the runner default. API keys go into the CI secret store, never into the config file.

Handling flaky and costly evals

LLM outputs are non-deterministic, and a gate that fails randomly gets muted within a week. Three controls keep it trustworthy. Set temperature to zero (or the provider equivalent) for deterministic assertion classes, and reserve sampling variance for tests where it is the point. Repeat each test a small number of times and require a minimum number of passes instead of a single verdict, which filters one-off sampling noise at a modest token cost. Split the suite by cadence: fast deterministic checks on every pull request, rubric-graded evaluations on demand or nightly, and full benchmark sweeps with the harness only on model changes. Log every run’s aggregate score to a time series; a slow two-month decline in pass rate is a regression your per-commit diff will never show.

Sources