# From Vibe Checks To Regression: Why LLMs Need Quality Gates In CI

The feature looks fine in the demo. The prompt is tidy, the answers are fluent, the JSON parses, nobody raises an eyebrow. A handful of runs later, it still behaves. Ship it. A week on, the same prompt returns malformed output three times in a row, just enough to break a downstream parser and wake someone up at 02:00. Nothing dramatic. Just failure that repeats often enough to matter.

This pattern keeps showing up because the initial check feels reasonable. Competent engineers look at an LLM feature and see something that doesn’t crash loudly. There’s no stack trace, no core dump, no obvious invariant to assert. You can’t assert truth, so teams assert plausibility. The output sounds right, the structure mostly holds, and the edge cases feel academic. Under time pressure, that feels like judgement. It rarely gets called negligence until later.

There’s also a cultural mismatch. CI grew up around deterministic code paths, fixed inputs, and outputs that either match or don’t. LLM features arrive sounding like interfaces, but behaving like stochastic services with a memory.

Where this breaks is not the demo path. It’s repetition, drift, and the quiet accumulation of small changes. Prompt edits that move a comma. A model version bump that nudges tokenisation. A temperature tweak from 0.2 to 0.4 because someone wanted “more creativity”. Run that fifty times and the failure mode appears.

The parser still works.

The semantics don’t.

The incident reports tend to read the same way. The output was “valid”, but different. A price field that used to be an integer becomes a string with a currency symbol. A refusal that was rare now triggers at a 3–5% rate for a previously green prompt. None of this trips alarms if the only check is “did it return something”.

This is where teams accidentally skip defining a **behavioral contract**. Not a formal spec. Just an implicit agreement. The contract is there anyway, enforced by consumers who assume that a field stays numeric, that a refusal only happens under certain prompts, that a denylist actually denies. When that contract shifts silently, production absorbs the cost.

The uncomfortable part is that LLM behaviour changes even when nothing “breaks”. Without baselines, teams don’t notice because there’s nothing to diff against. Yesterday’s 1% schema violation becomes today’s 6%, and it only surfaces when the nightly job chokes.

A defensible gate doesn’t try to make the model deterministic. That ship has sailed. It acknowledges variability and still draws lines. It asks what changes are acceptable and which ones should stop a release. That means choosing where in CI to be annoying. Blocking a PR is expensive. Letting everything through and watching dashboards is cheaper, until it isn’t.

### What a defensible gate would block on

**1\. Structural stability against a baseline**

Given a fixed prompt set and seed strategy, schema validity must stay within an agreed band compared to the last known good run. A jump from ~99% valid JSON to 94% is not “noise”; it’s a change that needs justification.

**2\. Semantic invariants under repetition**

Fields declared numeric remain numeric. Identifiers don’t quietly change format. Refusal rates for allowed prompts stay within a defined ceiling relative to baseline. If refusals move from 0.5% to 4%, the gate blocks, regardless of how polite the text sounds.

**3\. Safety guarantees as hard constraints**

PII denylist hits must remain zero, and redaction coverage must not regress below the previous release’s diff. Any new leak blocks, even if everything else improves.

This isn’t about personal habits. These are design requirements that force an uncomfortable conversation. Someone has to explain why a diff is acceptable. Someone has to decide whether speed beats control this time.

The mechanics are unglamorous. They also tend to be the first thing teams postpone. Warn on small deltas, block on large ones. Compare against stored outputs, not ideals. Track distributions, not single examples. None of this is free. It costs tokens, CI minutes, and attention. Blocking a release because a refusal rate crossed an arbitrary-looking line feels bureaucratic until the on-call rotation fills up.

There’s a real trade-off here, and pretending otherwise doesn’t help anyone. Tight gates slow teams down and occasionally block harmless changes. Loose gates keep velocity high and push risk downstream. Monitoring-only setups feel modern and flexible, but they assume someone is watching and empowered to stop things after the fact. CI gates are blunt, but they fail early, when fixes are cheaper and reputations aren’t involved.

What’s striking is how often teams accept silent behavioural drift as the price of using LLMs, while never accepting it in any other dependency. A payment library that changed number formats between patch versions would be rolled back in minutes. A model that does the same gets a shrug and a Slack thread.

This isn’t about building fortress walls around every LLM feature. It’s that “looks fine” is not a quality signal once repetition enters the picture. We keep shipping demos into systems that demand contracts, and then act surprised when the contract turns out to matter.

Some will argue that this overfits today’s tooling, that the gates will rot, that humans should review outputs instead. Maybe. Or maybe the discomfort is the point. The alternative is treating vibes as a regression strategy.
