agentic systems · production lab

Getting an agent you can trust: Amazon Bedrock AgentCore, one concern at a time

Spinning up a demo agent is a weekend. Building one you would let near a customer is the hard part: it has to run somewhere isolated, prove who it is, reach your systems without holding your passwords, remember the last conversation, and let you see and prove what it did. Nine sections, each with the plain-English version, every configuration option and what it decides, and simulations you can break.

The gap this exists to close

In plain English: getting an AI agent to work once, on your laptop, is a weekend. Getting one you would let near a customer is the hard part. It has to run somewhere safe, prove who it is, reach your systems without holding your passwords, remember the last conversation, and let you see and prove what it did. AgentCore is AWS selling you those seven concerns as managed pieces, so you assemble rather than build.

Cross the chasm: add one concern at a time

Try this: start with nothing and switch each piece on in order. Each one fixes a specific way the previous version was not production-ready. Turn one back off and the failure it prevented comes straight back.
in real lifeA prototype car that drives beautifully round a car park. Turning it into something you can sell means brakes, seatbelts, a number plate and an MOT, none of which made it go any faster.
start bystart with everything off, then switch on the pieces one at a time and read what each was protecting you from.
words hereprototype something that works on one machine, for one personproduction something a company can rely on, with everything that impliesisolation keeping one customer’s work away from another’s

Runtime: every configuration, and what each one decides

In plain English: Runtime is the serverless box the agent runs inside. You do not pick a server size. You pick a protocol (how callers talk to it), a network posture (what it can reach and who can reach it), a filesystem (what survives), and session timers (how long it may sit there). Those four choices are most of the design.

Session lifecycle: and what you are actually billed for

Try this: drag the timers and watch when the box is deleted. The billing line is the part people get wrong: CPU is charged only while the agent is actually working, not while it waits for the model to answer, but memory is charged the whole time the session is alive.
in real lifeA hotel room billed by the hour. Checking out takes it off the bill; leaving the key in the door does not, and nobody comes to tidy up until long after you have gone.
start bydrag the idle timer down and watch when the deletion actually happens.
words heresession one customer’s working spaceidle timeout how long it sits doing nothing before being cleared upbilled what you pay for, usually the time it existed, not the time it was used

Identity: two different doors, two different locks

Try it: count the integrations

Try this: drag the number of agents and tools. The figure that matters is not how many boxes there are. It is how many connections somebody has to build, authenticate and maintain.
in real lifePlugging appliances into sockets. Ten appliances and thirty sockets means three hundred possible cables if each one needs its own. One extension board means forty.
start bydrag tools up to 60 with the gateway unticked, and read the number.
words hereintegration one wired-up connection between an agent and a toolgateway one door everything goes through, instead of manycredential the secret needed to use a tooltool an ordinary function the model can choose to call

Gateway: one door to every tool

Policy, Registry, Browser, Code Interpreter, Payments

Try it: eight jobs, and the primitive each one actually needs

Try this: eight things somebody will ask an agent to do. Pick the primitive you would reach for, then see what the job really needs. Three of these are commonly given to the wrong one, and each wrong answer fails in a way that looks like the model being stupid.
in real lifeA toolbox where three of the tools look interchangeable. A screwdriver will turn a bolt, badly, once. The job still gets done, so nobody notices the wrong tool was used until the bolt is rounded off.
start byanswer read the latest figures off a supplier’s website, almost everyone picks the wrong one first.
words hereprimitive one of the managed building blocks the platform gives youtool an ordinary function the model can choose to callBrowser a real cloud browser the agent can click and readCode Interpreter a sandbox that runs code the agent wroteGateway one front door that turns existing systems into toolsMemory what the agent still knows in a later conversationsandbox an isolated space where code cannot reach anything elseAPI a defined way for one system to ask another for something

One conversation, four strategies: see what each keeps

Try this: the same three sentences go into every strategy. What comes out is completely different, which is the clearest way to understand why picking the wrong one leaves your agent unable to do the thing you wanted.
in real lifeFour people taking notes in the same meeting. One writes everything, one writes a summary, one writes only decisions, one writes only what is relevant to them. Later, they disagree about what was said.
start byread what each strategy kept, then look at what all four of them threw away.
words herememory what the agent still knows in a later conversationsummarise compressing what happened into fewer wordsextraction pulling out particular facts and keeping only those

Memory: what the agent keeps, and where you put it

Watch one request become a trace

Try this: send a request and watch the spans arrive in real time, exactly as they would in a trace waterfall. Then break something and see which span turns red. That is the whole value of observability in one picture. A non-deterministic system you cannot see is a system you can only apologise for.
in real lifeA parcel tracking page. Each scan tells you where it is and how long it sat there, and the delay you are complaining about turns out to be one depot, not the whole journey.
start bysend one request and watch which bar is widest.
words heretrace the record of one whole requestspan one step inside it, with its own start and end timelatency how long the user waitsJWT a signed token carrying who you are, which the receiver can verifytool an ordinary function the model can choose to call

Try it: why “200 OK” lies, and the plumbing that tells the truth

Try this: the request came back in 1.12 seconds with a clean 200, and every infrastructure dashboard is green. Turn the signals and the OpenTelemetry plumbing on one piece at a time, and watch the same request go from “healthy” to “it read a stale document and answered wrong.”
in real lifeA courier marks a parcel “delivered, 200 OK” and the tracking page turns green. It never mentions that the parcel was left at the wrong door. Fast and delivered is not the same as right.
start byleave everything on and read the backend panel. Then try one switch at a time: turn the collector off and back on, drop sampling to 1%, and switch on the evaluator span.
words hereOpenTelemetry a vendor-neutral standard for emitting traces, so any tool can read themsignal one kind of telemetry: a metric, a log, or a tracemetric a number counted over time (latency, error rate). Cheap and always on, but aggregate: blind to any single wrong answerlog a line of text a service writes as it runs; detailed, but hard to connect across servicestrace the record of one whole requestspan one step inside a trace, with its own start and endattribute a labelled fact attached to a span: which model, how many tokens, which documenttokens the pieces of text a model reads and writes; you are billed per tokentraceparent the header that carries the trace id downstream, so every service’s spans join the same tracecontext propagation passing that header on at every step; skip it and the trace shatters into fragmentsexporter the piece inside your app that batches spans and sends them onwardcollector the service that receives spans, processes them, and forwards them to a backendbackend where traces are stored and viewed (Cloud Trace, CloudWatch, Jaeger)sampling keeping only a fraction of traces to save cost; head sampling decides before the request finishes, tail sampling afterevaluator a check that scores the answer’s quality and emits its own spanground truth the known-correct answer you score againstHTTP 200 the success status code: the request completed. It says nothing about whether the answer was rightretriever the step that fetches documents for the modelstale out of date: a document that no longer reflects the truthorphan fragment a span with no parent trace to attach to

Would this change pass the gate?

Try this: you have changed the prompt. Scores move. Decide whether it ships, then notice that an average can rise while the thing you care about falls.
in real lifeA blood test with several numbers on it. One improved, one got worse, and whether that is good news is not something a single number can tell you.
start byaccept a change where one score went up and another went down, then read the verdict.
words hereeval a set of test cases used to score a changegate the rule that decides whether a change is allowed to shipregression something that used to work and now works less welltool an ordinary function the model can choose to calltrajectory the sequence of steps the agent actually took

Observability and Evaluations: see it, then prove it

What the anti-patterns actually cost, per request

Try this: start with everything wrong, then fix them one at a time and watch the token count and latency fall. These are not style preferences, each fix is measurable.
in real lifeA house with the heating on and every window open. Each open window seems minor. The bill is the sum of all of them.
start byfix the most expensive one first and watch the per-request figure fall.
words hereanti-pattern a common way of building something that quietly costs youper request the cost of one user asking one thingcontext everything sent to the model each time. The main thing you pay for

Multi-agent: four patterns, and choosing between them

Try this: pick what you care about most. The honest answer is often that you do not need multiple agents at all, but when you do, the priority you choose rules out as much as it rules in.
in real lifeOrganising a team. One person doing everything, a manager handing out work, specialists in a line, or everyone in a group chat. Each is right for some jobs and a disaster for others.
start bypick lowest latency and see which pattern wins, then pick easiest to debug.
words heremulti-agent several agents working on one job between themsupervisor one agent handing work to others and collecting the resultslatency how long the user waitshand-off passing work from one agent to another, along with what it needs

Six anti-patterns that quietly ruin a multi-agent system

Try it: run the loop until the pull request is green

Try this: a change is pushed. Step the loop forward and watch it close, and notice where it stops on its own, because a loop with no stopping condition is not automation, it is a bill.
in real lifeA dishwasher that reloads itself. Fine, until it decides the plates are not clean and starts again, and nothing tells it to stop.
start bypress next step seven times, then untick the cap and keep going.
words herepull request a proposed code change somebody reviews before it landspipeline the automation that runs when code changesround one full pass of test, report, fix, re-test

A worked case: autonomous UI QA as a closed CI/CD loop

What a Solutions Architect exam asks about all this

Try it: drill the scenarios instead of reading them

Try this: the same scenarios as below, one at a time, with the answer hidden. The alternatives are real answers to other scenarios on this page, so none of them is obviously silly, which is the point.
in real lifeFlashcards versus a revision guide. Reading the guide feels like learning and mostly is not. You recognise the answer when you see it, which is a different skill from producing it.
start bypress next scenario without reading the table below first. Recognising an answer and recalling one are not the same thing.
words herescenario a situation described the way an exam or an interviewer would put itdistractor a wrong option that is a real answer to something else