One pipeline. Pick a way production tests it.
One question, one AI, one answer, with people watching. Everything below this line is a problem the demo never had to survive.
Right in the hour of an inspection, the AI service starts refusing requests because too many arrived at once. At the same time the serverless parts — small programs that only wake up when needed — are slow on their first run. Three answers. Wait and retry, with a spread-out delay called backoff. A queue that keeps everything in order. And a fast lane kept warm for requests that cannot wait.
The AI is asked for a tidy list and sends back something half-finished, or wrapped in stray characters. A strict format check (a schema) catches most of it, a repair step rescues what it can, and a plain fixed rule handles the rest. One bad reply spoils one line of one report. It never reaches the work waiting behind it.
Someone hides instructions inside an uploaded document, and the AI follows them as if they came from you. That is called prompt injection. Three defences. Every piece of text carries a note saying where it came from, its provenance. The AI can only touch things on an approved list, an allow-list. And a guardrail blocks any action nobody asked for.
AI can only read so much at once. Fill that space with everything you have, in no particular order, and the answers quietly get worse the more you give it. The fix has four parts. A token budget, meaning a firm limit on how much text goes in. Sorting, so the most useful material goes first. Summarising older material as you go. And a deliberate decision about what gets dropped, rather than letting it fall off the end.
Yesterday's conversation bleeds into today's answer, and afterwards nobody can work out how either one happened. Each conversation is kept separate, everything it remembers has an expiry time (a TTL) that really does expire, and any conversation can be replayed to see exactly what the AI knew at the time.
A name, an NRIC, a salary figure ending up somewhere it should never go: into a log file, into a message to the AI, or onto a computer in another country. Personal details are blacked out before anything is sent, which is redaction. Each field has its own rules. And where data is allowed to live, its residency, is enforced at the door instead of assumed.
The bill turns up, and it is the same question asked ten thousand times. Send the easy work to a smaller, cheaper AI, keep a copy of answers that never change so you only pay once, and set a spending cap that stops things before the invoice does.
An inspector asks why one particular decision was made, four months ago. Every decision keeps a record — a decision trace — holding the rule it came from, the evidence behind it, what the AI suggested, and the name of the person who changed it if anyone did.