Every demo of "add an LLM feature to this old system" ends at the prompt working once, in a notebook, on a good example. The actual project starts after that, when the call has to live inside a job queue, a cron schedule, or a request handler that was written on the assumption every downstream call is fast, deterministic, and either succeeds or throws.
An LLM call is none of those three things. It's slow enough to blow past timeouts tuned for a database query. It's nondeterministic, so "retry the exact same call" can silently produce a different answer the second time. And it fails in new ways — a rate limit, a content filter, a response that's technically 200 OK but semantically empty — that the surrounding error handling was never written to recognize.
The unglamorous 80% of this kind of work is reconciling those two sets of assumptions:
- Idempotency has to be explicit. A legacy retry loop that assumes "calling this again is harmless" is a correctness bug the moment the thing it calls is a model instead of a database write. Every retry needs a key, and every handler needs to check it before re-running.
- Timeouts need a second opinion. The timeout that was correct for a SQL query is almost never correct for a model call, and raising it everywhere just moves the problem to whatever's waiting on the other end of that request.
- "It returned something" isn't the same as "it succeeded." Legacy error handling is built around exceptions. A model that returns a confident, well-formatted, wrong answer doesn't throw anything — so the check for correctness has to move from the transport layer into the application layer.
None of this is specific to any one model or framework. It's the same work as introducing any slow, unreliable dependency into a system that wasn't built to expect one — it's just that LLMs make it easy to forget that's what's happening, because the integration itself looks so simple.