The appeal of chaining prompts is that it reads like Unix pipes: the output of one call becomes the input of the next. That mental model is exactly what gets teams into trouble, because a shell pipe and a chain of LLM calls fail in opposite ways.

When grep gets malformed input, it errors loudly. When step two of a prompt chain returns malformed JSON, step three doesn't error — it does its best to answer the question anyway. The model downstream has no concept of "this input is broken," only "here is some text, I should respond to it." So instead of a crash you get a confident, plausible-sounding answer built on top of a silent failure one hop back.

A few things have made this tractable in practice:

  • Validate at every hop, not just the last one. Each step's output gets checked against a schema before it's allowed to become the next step's input. If it doesn't validate, that's a retry or a fallback, not a pass-through.
  • Treat the chain as a state machine with named states, not a sequence of function calls. Each transition has an explicit set of valid outcomes — success, retryable failure, terminal failure — rather than an implicit "it returned something."
  • Log the intermediate states, not just the final answer. The failures that matter are almost never visible in the final output. They're visible in the gap between what step two returned and what step three assumed it would receive.

None of this is novel if you've built distributed systems before — it's the same instinct that makes you put a circuit breaker between two services instead of trusting the happy path. The difference with LLM chains is that the failure mode is quieter: a broken service call throws, a broken prompt chain just writes a worse paragraph.