Writing
The boring part of shipping an LLM feature into old software
The hard part is never the prompt. It's retries, timeouts, and idempotency when a job queue built for deterministic calls now has to live with a model that isn't.
What VQA taught me about evaluation debt
A benchmark score is a hypothesis about your users' questions. Most teams ship before they've tested that hypothesis against a single real photo.
Prompt chaining is a state machine, not a pipeline
Treat step two's malformed output as a shell pipe and step three won't error — it'll hallucinate around the gap. Chains need guards, not just arrows.