Free estimate
Menu

Articles

LLM pipelines that stay cheap and correct in production

Key takeaways

  • Screen every item with a small model on short input; in our pipeline only about 8 % go on to full-document processing.
  • Two-stage routing sends the screening model about 5× less input than screening every item on its full documents.
  • Replay a stratified sample through the old and new designs before release, and look for items the new design would drop.
  • Validate output with a strict parser in code, make writes conditional on job state, and cap each grade by the evidence behind it.
  • Meter every call by stage and model, and reconcile those estimates with the provider’s per-model billing every day.

An LLM pipeline that works on a hundred items can cost too much on a million, and a small error rate becomes thousands of bad records. Cost grows with the tokens you send, so the design question is which items deserve which model and how much input. Correctness at volume depends on what the code does with each answer before it reaches the database.

We run LLM classification pipelines in our own production systems: about 1.4 million LLM processing runs so far. This guide covers the practices that keep that pipeline affordable and its output trustworthy: routing by cost, replay tests, strict parsing, conditional writes, evidence caps, metering, and price-floor routing.

Screen every item with a small model on short input

Most items in a classification pipeline do not need the expensive path. Screen each one with a small, cheap model on short input, such as a title and a description, and let only the items that pass go further.

Our pipeline routes by cost. A small model screens each item on short input. Only about 8 % go on to full-document processing, and only items that pass the screen get long-form output from a premium model.

# Illustrative.
def route(item, screen, deep, report):
    """Two-stage routing: cheap screen first, full documents only when needed."""
    verdict = screen(item.title, item.summary)        # small model, short input
    if not verdict.worth_reading:
        return verdict                                # most items stop here
    detail = deep(item.title, item.documents)         # full documents
    if detail.passes:
        detail.report = report(item, detail)          # premium model, long form
    return detail

Tune the screen toward recall. A false pass costs one full-document read; a false stop loses the item for good. Set the threshold so the deep stage sees every plausible item, and let the deep stage do the precise work.

The screen needs a clear question with a cheap answer. “Is there enough overlap to read the documents?” works. “Write a summary” does not belong at this stage, because long output costs as much as long input and teaches the router nothing.

Compare the design with and without routing

State the saving as a comparison between two designs on the same model, so the number means something. Two-stage routing sends the screening model about 5× less input than screening every item on its full documents. Both sides of that comparison use the same small model; the difference is only what each item sends.

Compare like with like when you report a ratio of this kind. A saving measured against “send everything to the large model” mixes two changes, a smaller model and less input, and nobody can tell which one paid. One change per comparison keeps the number honest and repeatable.

The premium model is the third step, not the second. It writes long-form output only for items that passed the screen and the full-document read. Its cost then follows the items that matter, not the items that arrive.

Replay a stratified sample before release

A routing change can save money by quietly dropping items that belong on the deep path. The test for that is a replay: run a fixed sample through the current design and the candidate, and compare the decisions item by item.

Stratify the sample by outcome. A random sample of a stream where most items stop at the screen holds very few of the items you care about. Draw fixed numbers from each outcome class, and keep a same-day copy of the current design as the baseline, so the two runs see the same inputs.

# Illustrative.
from collections import Counter

def compare(sample, baseline, candidate):
    """Item-by-item comparison of two pipeline designs on one fixed sample."""
    moves = Counter()
    dropped = []
    for item in sample:
        old, new = baseline(item), candidate(item)
        moves[(item.stratum, old.grade, new.grade)] += 1
        if old.grade > 0 and new.grade == 0:
            dropped.append(item.id)    # items the candidate would lose
    return moves, dropped

Read the result in two parts. The moves table shows how grades shift between designs in each stratum, which tells you whether the candidate is stricter or looser overall. The dropped list is the one to read item by item: each entry is an item the current design kept and the candidate would discard.

We replay-tested the two-stage design, and a later model change, against stratified samples before release. In the first replay, 94 % of non-qualifying items stopped at the small model, and none of the rest was promoted. When we last changed models, candidate replacements were tested by replay before any went into production.

Parse model output strictly, in code

Treat model output as untrusted input. Parse it with a strict parser in code, check every field the database needs, and reject anything that does not fit. Let a rejected answer fail its job so it runs again, rather than write a partial record.

Keep any repair step mechanical: strip code fences, normalize whitespace, round a fractional grade to the scale. A model asked to repair another model’s output may change the content as well as the format. The repaired record can then say something the first answer never said.

# Illustrative.
import json

GRADES = {0, 1, 2, 3, 4, 5}        # illustrative scale

def parse_report(raw: str) -> dict:
    text = raw.strip().removeprefix("```json").removesuffix("```").strip()
    data = json.loads(text)                       # raises on malformed JSON
    grade = round(float(data["grade"]))
    if grade not in GRADES:
        raise ValueError(f"grade out of range: {grade}")
    body = data.get("report", "")
    if grade >= 4 and not body.strip():
        raise ValueError("a high grade needs a report")
    return {"grade": grade, "report": body}

Make every write conditional on job state

A job can be claimed twice: a lease expires, a worker restarts, a retry overlaps a slow first attempt. If both runs write, the last one wins, whatever its quality. Make each write conditional on the state the job was in when this run claimed it.

-- Illustrative.
update jobs
set    grade      = $2,
       report     = $3,
       status     = 'complete',
       updated_at = now()
where  id = $1
  and  status = 'processing'
returning id;

If the update returns no row, another run has already finished the job, and this run discards its result. Writes in our pipeline are conditional on job state in this way, so a stale duplicate cannot overwrite a finished record.

Cap each grade by the evidence behind it

A model will grade an item on thin input as confidently as on rich input. Bound the grade in code by the amount of evidence the model actually saw, so a high grade always rests on enough material to justify it.

# Illustrative.
def bounded_grade(grade: int, evidence_chars: int, has_documents: bool,
                  min_chars: int, cap_when_thin: int) -> int:
    """Cap a grade when the model saw too little to support it."""
    if not has_documents and evidence_chars < min_chars:
        return min(grade, cap_when_thin)
    return grade

Set the thresholds from your own data, and keep them in code beside the parser rather than in the prompt alone. A rule in the prompt is a request; a rule in code always holds. In our pipeline, output grades are bounded by the amount of input evidence.

Meter every call by stage and model

You cannot control a cost you only see as a monthly total. Record tokens for every call, tagged with the pipeline stage and the model, so a change in spend points at its cause.

Metering does not have to touch the production workflows. Our workflow engine already keeps a record of each execution, and a harvester reads token usage from those records by stage and model. A daily job then reconciles our estimates with the provider’s authoritative per-model billing.

-- Illustrative.
-- Daily reconciliation: our metered estimate against the provider's bill.
select u.day, u.model,
       sum(u.estimated_usd)                         as metered,
       b.billed_usd                                 as billed,
       b.billed_usd - sum(u.estimated_usd)          as gap
from   token_usage u
join   billing_daily b using (day, model)
group  by u.day, u.model, b.billed_usd
order  by u.day desc, abs(b.billed_usd - sum(u.estimated_usd)) desc;

A steady gap means the price table is out of date. A gap on one model on one day means a route or a provider changed, and that is worth a look the same day.

Pin bulk work to price-floor providers

Model aggregators can serve one model from several providers at different prices. Left alone, a router may send bulk traffic to a provider that costs more than the model’s list price. For bulk workloads, pin the route to the providers at the price floor.

We pin bulk LLM workloads to price-floor providers and check billed cost against list price, per model. The daily reconciliation above is the check: a model whose billed cost drifts above list shows up in the gap column.

Treat a cheaper model as a release, not a setting. Replay it against the same stratified sample as any other change, and switch only when its decisions match. Watch its error rate and latency under real traffic as well, because a model that is cheap per token can still cost more per finished item.

Settings we use

SettingValueWhy
Screening inputShort input only, on a small modelMost items stop here, so the cheapest call does most of the work
Full documentsOnly items that pass the screen; about 8 %About 5× less screening input than screening every item on its full documents
Long-form outputA premium model, only for items that passThe expensive model’s cost follows the items that matter
Release checkStratified replay before releaseCatches items a change would drop; we tested the two-stage design and our most recent model change this way
Output validationStrict parser in codeMalformed output is caught before anything is written
WritesConditional on job stateA stale duplicate run cannot overwrite a finished record
Grade boundsCapped by the amount of input evidenceA high grade always rests on enough material
MeteringPer call, by stage and model, reconciled daily with billingSpend changes point at their cause
Provider routingPrice-floor providers for bulk workBilled cost stays at list price

Recommendations

  • Screen every item on short input with a small model, and send only the items that pass to full documents and premium models.
  • Report savings as a with-and-without comparison on the same model, so each number measures one change.
  • Replay a stratified sample through the current and candidate designs before every routing or model release, and read the list of dropped items.
  • Validate output with a strict parser in code and keep repair mechanical; never let a model rewrite another model’s answer.
  • Make every write conditional on job state, and cap grades by evidence in code rather than in the prompt.
  • Meter tokens by stage and model from the start, and reconcile them with the provider’s billing every day.

References

If you need an LLM pipeline that stays affordable at volume, see Infrastructure and data or get a free estimate.

Send us the system that has to stay fast and affordable.

Get a free estimate