the whole guide, distilled

Cheatsheet: the whole guide on one page

Every lab’s must-remember points, on one page you can scan the night before an interview or print and mark up. Each card links to the lab it came from, where the same points arrive with a working simulation behind them. Read it to scan, or press Test yourself to hide the points and recall each lab’s from its name, clicking any point to check. Use the filter to jump to a word: checkpoint, drift, KV cache, catastrophic, reducer.

Mental modelsthe shapes worth memorising, in technical form
residual stream + + LayerNorm → Multi-Head Attn Q,K,V · softmax(QKᵀ/√d)V LayerNorm → MLP up 4d · GELU · down
Transformer block: pre-norm, attention then MLP, each written back to the residual stream. Depth stacks this N times.
RunnableParallel context: retriever|rerank question: Passthrough prompt model parser {context, question} → prompt | model | parser
LCEL: the pipe is RunnableSequence; a RunnableParallel keeps the question alongside retrieved context, else the slot arrives empty.
work / token context length no cache: O(n) per step KV cache: O(1) per step
KV cache: keys and values depend only on token and position, so cache them. Decode goes from quadratic total to linear, trading compute for memory.
Q Kᵀ Q Kᵀ / √dₖ softmax × V attention(Q,K,V) = softmax(QKᵀ/√dₖ) V
Scaled dot-product attention: scores are query-key dot products scaled by 1/√dₖ, softmaxed into weights, then a weighted sum of values.
offline evalper commitgroundedness, task success (golden set) online proxyper releasethumbs-down, retries, escalation, reopen business KPIper quarterCSAT, deflection, cost per resolution validate the joints, do not assume them
The metric chain: three layers at three cadences. Each exists because the one above is unmeasurable that often; gate releases on the joints.
predicted actual positivenegative posneg TP FN FP TN precision = TP/(TP+FP) · recall = TP/(TP+FN) Fβ weights recall β² times precision
Confusion matrix: precision guards false alarms (the FP column), recall guards misses (the FN row). Report the pair; a single F score hides the trade.
modelproposes your codevalidate + gate toolruns {name,args} result back into context, matched by call id
Tool calling: the model emits a structured call and executes nothing; your code validates identity, args and policy, runs it, and feeds the result back. The boundary is yours.
node A node B reducermerge state ckpt crash resumes from the last checkpoint
LangGraph super-step: nodes due now run, their partial updates merge through reducers, then a checkpoint is written. A crash resumes there, not from the start.
Wd×d frozen + B d×r · A r×d train B,A only · r ≪ d · <1% of params
LoRA: freeze W, learn a low-rank correction BA with rank r far below d. Under a percent of the weights train, so optimiser state shrinks by the same factor.
cut where cumulative ≥ p nucleus (kept)tail (dropped)
Top-p (nucleus): sort tokens, keep the smallest set whose probabilities reach p, renormalise, sample. The set size adapts to confidence; top-k does not.
FP1616 bit · 2× base INT88 bit · half the memory INT44 bit · a quarter
Quantisation: fewer bits per weight, proportionally less memory and bandwidth. Quality holds down to about 4 bits, then falls off sharply; measure, do not assume.
Roofline (A100 FP16: 312 TFLOP/s, 2.0 TB/s, ridge 156 FLOP/byte): below the ridge you are memory-bound and faster maths changes nothing; only fewer bytes per FLOP helps.
Build, run and govern an agent

ADK

why an agent picks the tools it does

  • Wrong tool chosen despite the right one existing: blame the docstring, not the model; the model picks from generated schemas and a vague description makes two tools indistinguishable.
  • instruction is read by this agent’s model every turn; description is read by other agents and by Enterprise routing. Swap them and the agent works perfectly and is never used.
  • An agent is four parts: definition (you write), model (you pick), tools (you write), runner (ADK gives you). You own three of four.
  • output_schema (forced JSON) cannot combine with tools or delegation; if you need both, split it: one agent acts, a second formats.
  • Four state scopes by key prefix: none (this conversation), user: (one person everywhere), app: (every user), temp: (this turn only). User data in app: is a data breach.
  • State that vanishes after restart: something wrote to session.state directly instead of via tool_context.state or output_key, so the change never became a traced event.
  • Model call silently skipped: a before_model_callback returned a value; anything other than None short-circuits what it wraps.
  • Model-written code runs in an Agent Sandbox: one per session, no network, deleted in a finally; the runtime holds the credentials, the sandbox never sees them.
  • Deployed and healthy but no traffic from Enterprise: its description is too vague to route on. Nothing is broken; nothing is being sent.
  • A skill that silently never loads: the folder name does not match the name in its frontmatter, and there is no error for this.

Pick your frame

one loop, five philosophies

  • Every framework wraps the same loop (situation, model, tool, back); the product is an opinion about who owns the next step: a graph (LangGraph), a crew (CrewAI), the model (Strands), a kit (ADK), or the agent itself (Hermes-style).
  • Why not a while-loop: for one or two tools, write it yourself; adopt a frame only when sessions, retries, durability and evaluation cost more to hand-roll than to adopt.
  • LangGraph vs CrewAI: a knowable process that must survive crashes and long human waits wants LangGraph (drawn graph, typed state, checkpoints); work that maps to human job titles wants CrewAI (prose roles).
  • LangGraph writes a checkpoint every super-step; a crash resumes from the last one, so a multi-day approval is just a long gap between checkpoints.
  • Two parallel LangGraph nodes writing one state key errors on purpose; the fix is a reducer on that key (append to a list), a data-model decision, not “avoid parallelism”.
  • Model-first (Strands) gives up predictability: least orchestration code, most run-to-run variance; its partners are tracing and eval gates, not more orchestration.
  • CrewAI hierarchical adds a manager that catches gaps but bills the coordination (about 28% of a toy run); its own fix for mostly-fixed processes is Flows (deterministic rails in code).
  • Self-editing (Hermes) breaks review-then-release: the artefact changes without a deploy, so ops becomes continuous eval, versioned history and rollback; give autonomy in proportion to blast radius.
  • The Lang family by layer: LangChain builds, LangGraph orchestrates, LangSmith and Langfuse both observe, Langflow draws. Langflow and Langfuse are NOT the LangChain company.
  • A Runnable shares one interface (invoke, batch, stream); A | B is sugar for RunnableSequence. A RAG chain needs RunnablePassthrough to carry the question, because retrieval consumes the input.
  • RunnableParallel runs every branch on the same input (“do all of these”); RunnableBranch runs the first matching one (“do whichever fits”); RunnableLambda wraps a plain function.

AgentCore

an agent you would trust in production

  • AgentCore is AWS selling seven production concerns as managed pieces (run safely, prove identity, reach systems without holding passwords, remember, observe); you assemble rather than build.
  • Session isolation: each session gets its own microVM, so sessions cannot read each other’s files; ordinary serverless shares a sandbox.
  • The serverless runtime caps at 8 hours per session; for GPUs or longer use Runtime instances (managed EC2), up to 14 days.
  • Gateway is one door to every tool: 10 agents and 30 tools hand-wired is 300 integrations; one gateway makes it 40, and holds the credential so the agent never sees the downstream secret.
  • Hundreds of tools degrade model choice: 300+ tool descriptions per request costs accuracy and tokens, so use Gateway tool search (listing 360 tools can cost more than the call itself).
  • Identity means no long-lived credentials: OIDC federation into a role, a token vault for outbound, three-legged OAuth to act for a user (the user consents, the agent never holds the secret).
  • Short-term memory dies with the session; personalisation across sessions needs long-term memory (user-preferences strategy). Asking for one when you meant the other is the usual mistake.
  • Observability answers what it did; evaluation answers whether it was any good. A non-deterministic system needs both, from day one, because retrofitting them is far harder.
  • Prove tools ran in the right order with trajectory evaluators (expected_trajectory), programmatic not a judge; order matters for audit and security.
  • Policy is evaluated and enforced before the tool call; a prompt instruction is only advisory, so enforce sequences with an orchestrator, not a request. Context is a budget: attention is U-shaped, so quality degrades before the model’s limit.

Agents building agents

spec to running service

  • Spec-driven development: you describe the behaviour and a coding agent writes it, so the valuable artefact is the description; a good spec regenerates the whole codebase.
  • Two starting points: build-then-secure, or start locked-down and add what is needed. The second (infrastructure-first) answers “where does our data go?” before anything ships.
  • Gherkin (Given, When, Then) forces the starting state, the trigger and the outcome; prose lets you write “handled appropriately” and feel finished. The arguing moves from the spec (cheap) to code review (expensive).
  • Out of the box the coding agent targets the wrong API, usually the older style that cannot do graph routing or session state; nothing errors, it just quietly is not what you asked for.
  • A session-start pattern anchors the agent on the current API surface up front, because long conversations dilute context and drift back to deprecated models.
  • Ambient agents are triggered by an event, not chatted to; test by posting a fake event yourself, exactly as the broker would.
  • LLM-as-judge: output differs every run and you care about behaviour, so a model grades the whole trace against a rubric; string matching is brittle against probabilistic output.
  • Deterministic rule plus a judgement call maps to a graph: function nodes for the rule, a model node for the judgement. A dollar threshold belongs in Python; asking the model to do arithmetic invites the run where it fails.
  • Security in one sentence: your personal credentials never enter the sandbox where the agent runs code. Egress starts closed; hardening is mostly replacing the wide-open dev allowlist with the specific domains needed.
  • Every agent run is background: it returns an id and you poll until complete; the sandbox persists 7 days and passing its id back keeps files across turns.

Evaluation

how do you know it got better?

  • Traditional testing fails on generative systems: non-deterministic output means a binary pass/fail has nothing stable to assert. A test asks “did it pass?”; an agent answers “how often, and how badly, does it fail”.
  • Do not trust public benchmark scores: data leaks into training sets, and they over-weight tasks that do not resemble yours.
  • Three ways to score, none free: a ruler (cheap, only measurable things), people (the real standard, unaffordable at scale), or another model (autorater); an uncalibrated autorater is confidently wrong and you cannot tell from the numbers.
  • BLEU/ROUGE are poor for open-ended text: the single-ground-truth trap penalises valid paraphrase and rewards nonsense that reuses the reference vocabulary.
  • Absolute ratings measure inherent quality and stay reusable; side-by-side only says which of two was preferred and is sensitive to order and position.
  • Inspect the trajectory, not just the final output: the outcome says it failed, the trajectory says where and why, the only version you can act on.
  • Exact-match golden paths invent false negatives on any task with multiple valid routes; keep them as a narrow regression tool for tool schemas and argument parsing.
  • Tool calls per trajectory reveals a stuck agent: a spike means it is calling the same tool without progress.
  • Empty online-monitor results are a plumbing problem (telemetry not arriving, filter too narrow), never evidence of quality; keep monitoring affordable with a sampling cap.
  • Scores up on the eval set but production down means overfitting to the eval set or a local peak; that is what the held-out set exists for.
  • Use P99 not average latency: the average hides the slowest 1%, and those are the users who notice.

The number to trust

metrics for asymmetric mistakes

  • 94% accuracy is a trap on lopsided classes: “nobody needs help” scores 80% while missing everyone who did; accuracy mostly measures the size of the bigger group.
  • First answer to “which error is worse” is a question: what does each error cost? Report precision and recall separately, and recall leads when a miss is expensive.
  • When a miss costs more than a false alarm, offer F2 (a miss counts four times a false alarm), never F1, and still show the precision/recall pair underneath.
  • Precision 0.60/recall 0.75 and precision 1.00/recall 0.50 give the same F1 to three decimals, yet the second misses twice as many people; track the pair.
  • Same mean absolute error is not equally good on an ordinal scale: two one-band slips equal one two-band catastrophe; use quadratic weighting or the full band table.
  • Trust an LLM judge only by measured agreement with humans on a labelled sample; below about 0.6 agreement every rubric score inherits the judge’s noise. Pin its version; never let the generator judge its own output.
  • Fluency is not grounding: if the right passage sits below the context cutoff the report is confidently grounded in the wrong evidence. Recall@k is the ceiling on everything downstream.
  • Distribution drift needs no labels and no change on your side; a PSI check against a frozen baseline is the cheapest alarm there is, and the gauge trips at PSI 0.25.
  • Never leave a metric unsplit: overall recall 0.62 can hide per-slice 0.75, 0.75 and 0.40; slices are the fairness check you can run on data you already hold.
  • Sample size: with 8 labelled examples the 95% margin is about plus or minus 35 points, so 0.82 honestly means “between half and perfect”. A metric without a sample size is a vibe with a decimal point.
  • Deterministic gates first, judgement later: run the checks that cannot flake (100 runs, 100 identical verdicts) before anything reaches a person.

Govern and secure

gateway to audit trail

  • Set boundaries by the worst action a hallucination could cause, not “what does it need” (which only grows). If it must not happen, remove the capability; instructions are soft controls that compete with injected text.
  • Indirect prompt injection hides a malicious instruction in content the agent fetches (a retrieved document), so the attacker never interacts with your system. This is the threat that makes agents genuinely different.
  • Model Armor evaluates text during the CONTENT_AUTHZ phase (both request and response bodies), before it reaches the reasoning engine; REQUEST_AUTHZ only reads headers.
  • The egress path routes through a dedicated network attachment in your own network; it controls the agent’s ability to call out and is what prevents data exfiltration.
  • Each agent gets a unique, rotated, short-lived certificate identity per deployment, auto-revoked, so no static key to steal and an audit trail naming one agent, not a shared account.
  • High agent-to-tool 403 denials mean the identity lacks the role to reach that server, not a content block and not a registry problem.
  • Delegated (three-legged) identity lets the target enforce its own per-user access; a machine identity would grant everyone the same maximal view.
  • Service perimeters block API calls moving data outside the perimeter regardless of permissions; a principal access boundary is a second ceiling that still applies when IAM is misconfigured.
  • A policy enabled but blocking nothing: check the enforcement mode; a dry-run policy records what it would have blocked without blocking it.
  • The Agent Registry decouples agent code from infrastructure: update a tool’s address without touching agent code. Restrict an agent to read-only tools with an IAM condition in CEL.
  • Red teaming is a coverage exercise: the board of threat families times surfaces is the artefact, and every break found becomes a regression test forever.

Gemini Enterprise

rolling it out to a whole company

  • Agents reason, decompose tasks, run multi-step workflows and use tools to reach enterprise systems; a chatbot just answers.
  • Your data sources choose your identity provider: Workspace needs Google Identity; Microsoft 365 needs Workforce Identity Federation with Entra ID; people use Workforce, pipelines use Workload Identity Federation.
  • Mobile requires OIDC (SAML cannot serve it); OIDC is lower maintenance (Google polls the provider’s discovery URL), while SAML needs a certificate uploaded before the old one expires. Legacy ADFS offering a metadata XML = SAML.
  • Human supervision is guaranteed by a structured draft the user must explicitly authorise before execution, not an audit trail reviewed afterwards.
  • Raise backend timeouts to three to five minutes for generative AI; models routinely exceed the standard thirty seconds.
  • Use ingestion over federation when no connector exists; federation’s quality and latency depend on the source API, so it is discouraged for day-one value.
  • Cloud Storage and BigQuery connectors let all users access all connected data unless access-control lists are enabled.
  • Model Armor sits between user and model, sanitising both prompt and response; it covers prompt injection, sensitive-data disclosure, malicious files and unsafe URLs, not training-data poisoning, model theft or supply chain.
  • Floor settings are org/folder/project minimums; a template below them saves without error and defaults up to the floor. Disable prompt-injection detection on the output template, because those attacks originate in the prompt.
  • Check Grounding flags ungrounded statements but does not rewrite them; the Ranking API re-orders by relevance, and NDCG = 1 means the ordering could not be better.
  • Search tuning logs relevant and irrelevant answers to move the query toward the right content, on unstructured stores only; it adapts the query’s position, whereas fine-tuning adapts every embedding.

The FDE bootcamp, day one

same model, different body

  • Seven organs, seven places to attach a control: reasoning core, capabilities, context plane, runtime, orchestration, control plane, identity and data boundary. Safety is never a property of the model.
  • The session is the log: a tool write is a delta on the event, committed only on append. That buys audit, resume and consistency; crash before the append and the tool runs again, so tools must be idempotent.
  • State scopes: plain key for this session, user: across a user’s sessions (isolates tenants), app: for everyone, temp: for this turn and never persisted (where secrets go).
  • Compose: transfer when judgement should route and the specialist then owns the chat; agent-as-tool to keep the caller in control and see only the result; a graph workflow for deterministic, auditable flow. A join node is a barrier; a plain fan-in node fires per branch.
  • A remote agent has no output key: add an after-agent callback that copies its reply into state, or its branch is empty at the join.
  • MCP inside agents, A2A between agents. A capability is MCP; delegating to something that reasons for itself and is owned by someone else is A2A.
  • The agent card names identity, skills, capabilities, interfaces (each with a protocol version) and security schemes; identity rides in transport headers, never the card or the body. Advertise 0.3 beside 1.0 or the enterprise app will not register it; the console custom-agent flow bypasses the gateway.
  • 401 versus 403: no credential versus a known identity without the grant. Locally the playground calls as you; deployed, as the runtime service account. Private Cloud Run wants an identity token with the service URL as audience plus the invoker role; a redeploy resets that binding.
  • The lethal trifecta: private data, untrusted content, a way out. Screen inputs, scope reads to the asking user, guard egress with a before-tool hook, approve outbound actions; audit tells you what happened but stops nothing.
  • Delegation names both parties (user in sub, agent in act); impersonation erases the actor. A safe token is short-lived, delegated, holder-bound, one audience, narrow scope. Audience binding and no passthrough break the confused-deputy chain.
  • Own authority (workload identity, least privilege) for jobs that belong to the agent; on behalf of (scoped, short-lived delegation) whenever the action is a person’s. Hop-level authorisation at the gateway is necessary and not sufficient; rows are decided at the source.
  • A discovery engine evolves candidates against an evaluator you write; a population with a map of elites escapes the local dip a greedy climber stalls in.

Peak Week

one agent, all the way through

  • One agent is carried through six ordered stages: Groundwork, Make it work, Make it survive, Make it safe, Make it good, Make it theirs; each quietly leaves behind something a later one needs.
  • First fix for an agent grounded in testing but inventing in production is a graded eval with a reference-free grounding judge; you cannot fix a hallucination rate you have not measured, and invented sentences read exactly like grounded ones.
  • Every enterprise-portal chat failing while CLI works = the runtime is not normalising the session id: Enterprise passes a full resource path and the session service rejects an id containing a slash.
  • A deployed agent returning an empty reply after tool calls is an exception inside it; pull reasoning-engine stdout/stderr from Cloud Logging, the traceback is nowhere else.
  • Deploying to the managed runtime registers the agent and gives it an identity, but does not narrow its permissions; least privilege is the audit you still owe.
  • Memory Bank empty but recall works across sessions = the recall came from conversation history and the write side is missing (an after_agent_callback that adds the session to memory).
  • “Event loop is closed” on the second request = an async client created once and reused, bound to a closed loop; fix is a fresh client per request. A bug sparing the first request survives casual testing.
  • Bundle the data inside the deployment instead of reading from a bucket: it removes a permission and a failure mode at once and needs no storage role.
  • A tool call slow only on the first request is a cold sandbox warming up; re-send it, latency on a warm-up path is not a failure.
  • An agent flagging every run during a station-wide failure is reasoning from the depot rather than reconciling per service; the join is on leg_id not depot, and false positives cost a real journey.
Cloud and architecture

The hot seat

the eight probes that decide the room

  • RAG for ten thousand documents: parse and OCR, clean, chunk with overlap, embed, index, retrieve top-k, re-rank, generate with citations; measure retrieval recall separately from answer quality.
  • The bill is three times too high: requests times tokens times price, so measure first, then cache the repeated prefix, route easy requests to a small model, trim context, cap output, compress wording last; report tokens per request before and after.
  • “Ignore previous instructions”: assume the model will sometimes obey and make obeying harmless; data versus instructions, least-privilege tools, input and output guardrails, logging, a person for anything destructive.
  • Prompt, retrieval or fine-tune: volume, rate of change, knowledge versus behaviour, latency and cost budget. Fine-tuning teaches style and format, not facts that changed yesterday.
  • Model deprecated: eval set first, pin versions, shadow run, diff regressions, fix prompts, canary, cut over, keep rollback. Never “fine-tune the new model to match the old”.
  • Temperature reshapes next-word probabilities (lower is more predictable); top-p keeps the smallest set of words whose probabilities add up to p. Higher temperature is not more consistent.
  • Workflow versus agent: code fixes the steps in advance versus the model picking the next step at run time. The follow-up is state: where progress lives, who may write it, how a crashed run resumes.
  • One more for cheaper means a lever on the bill (cache, tier, trim, cap, batch), not a development saving.
  • Weak answers name categories; strong answers name parts, numbers and failures. Specificity is the difference, not length.

How the cloud works

request journey, OSI, landing zones

  • OSI is a fault-bisector, not seven steps: at 3am work up the stack, is there a route, does the security group open, can they speak the language, did they understand.
  • What makes a subnet public: a route to the internet gateway in its route table, nothing else. The subnet gives the route, the public IP gives reachability; you need both.
  • A private-subnet database reachable from the internet: public accessibility is on, and it overrides the subnet, the route and the security group.
  • Reserved addresses: AWS/Azure take 5 per subnet, GCP takes 4; a /28 gives 11 usable, not 16, and on an autoscaling tier that shortfall arrives as an outage. Use 10/8, 172.16/12, 192.168/16 only.
  • You cannot resize a subnet in use; add a new one from spare space and migrate, which is why a good plan leaves blocks unallocated.
  • Connection times out with no error = a network control (missing route, security group, stateless ACL); permissions never hang, they answer at once. A network ACL is stateless, so the reply needs outbound 1024-65535 allowed.
  • AccessDenied while the role has the permission = an explicit deny above it: an SCP, a permissions boundary, an endpoint policy, or a KMS key policy.
  • Machine identity: never an access key in config; the code asks its surroundings for a short-lived token that expires within the hour.
  • SAP-C02 triggers: “admin access, no inbound ports, full audit” = Session Manager / Bastion / IAP; “forty VPCs, least overhead” = transit gateway / WAN hub (a hub needs 40 links, peering needs 780); “partner API, overlapping addresses” = PrivateLink (overlap is irrelevant because the networks never join).
  • “Private endpoint created but traffic still goes over the internet” = the private DNS zone is not associated with the VPC, so the name still resolves to the public address.
  • GCP is different: the VPC is global with regional subnets, so crossing a region needs no peering; on AWS/Azure the network is regional and multi-region multiplies the networks.

The three-cloud atlas

same jobs, three dictionaries

  • Same job, three names; the small print is the product. Learn one cloud deeply; the second is a delta list, not a second degree.
  • Naming habits: “Amazon X” is something you use, “AWS X” manages AWS itself; Google prefers a plain noun that says the job; “Microsoft X” reaches beyond Azure into 365 and other clouds.
  • Relational at global scale: Spanner on Google Cloud, Aurora DSQL on AWS since 2025. Azure’s multi-region-write answer is Cosmos DB, which is not relational.
  • Warehouse philosophy: BigQuery is serverless and bills per query scanned; Redshift is capacity you size (with a serverless option); Azure’s answer is the Fabric warehouse.
  • Kafka clients with no code change = Azure Event Hubs, which speaks the Kafka protocol; MSK and Managed Service for Apache Kafka are the managed-Kafka answers; Kinesis and Pub/Sub are their own protocols.
  • Archive latency changes DR design: Cloud Storage Archive and S3 Glacier Instant Retrieval answer in milliseconds; Azure’s Archive tier and Glacier Deep Archive rehydrate over hours.
  • Guardrails: AWS service control policies only deny; Google Organization Policy constrains what may be created; Azure Policy can also audit and auto-remediate.
  • A perimeter you can buy: VPC Service Controls on Google Cloud and Network Security Perimeter on Azure; on AWS you assemble it from endpoint policies and service control policies.
  • The agent layer: Bedrock AgentCore, Agent Runtime plus Agent Gateway inside the Gemini Enterprise Agent Platform, Foundry Agent Service. Prompt screening: Bedrock Guardrails, Model Armor, AI Content Safety.
  • Names rot: Entra ID (was Azure Active Directory), Microsoft Foundry (was Azure AI Foundry), Cloud Run functions (was Cloud Functions), Valkey inside ElastiCache and Memorystore. Check the atlas’s “Renamed lately” tab before saying a name in the room.
  • Breadth is a clue, not a verdict: the depth map counts what each provider lists per family; the rows carry the judgement.

Diagram lab

a picture that argues

  • A diagram is an argument aimed at somebody, not a description; change the audience and you should draw a different picture.
  • One diagram for four audiences becomes sixty boxes that answer nobody; four small diagrams is four arguments, not duplication.
  • Put a question at the top, not a title: “how does a payment reach the ledger, and what happens when the ledger is down?” It is the only reliable way to decide what to leave out.
  • One diagram, one level of zoom. C4: Context always, Container the workhorse, Component only for the genuinely complex, Code almost never; the Deployment view is where cloud service names belong.
  • Seven parts make a diagram trustworthy: the question, owner and contact, date and version, where it lives, a legend, boundaries, and a “not shown” note. Undated means unknowable.
  • Accuracy is almost never the problem; usability is. A diagram can be entirely accurate and still fail every reader.
  • An arrow is the most overloaded symbol in engineering; pick one meaning. Best: the arrow points the way the request goes, the label says what is carried.
  • Solid = synchronous, dashed = asynchronous is the highest-value distinction: it shows which failures cascade and which a queue absorbs.
  • Draw the failure paths, not just the happy path: mark retries, timeouts, fallbacks and the dead-letter route; nobody opens a diagram to learn how the system works when it works.
  • Non-functional requirements (availability, latency, cost, residency) drive the architecture; answer them before drawing, and keep the numbers in a table beside the picture.
  • Whiteboard interview: spend the first ten minutes not drawing, restate the problem, ask the scale question, ask what must not fail versus what may be slow, say what you exclude, then draw. Not asking about scale is the mistake.
Under the hood

GPU lab

the machine underneath

  • A CPU makes one instruction stream fast; a GPU spends silicon on arithmetic and hides slow memory by having other work to switch to. Reference A100: 108 SMs, 2.0 TB/s HBM, 312 FP16 tensor TFLOP/s.
  • Divide 312 TFLOP/s by 2.0 TB/s: the card wants roughly 150 operations per byte fetched; do less and the arithmetic units idle.
  • A warp is 32 threads in lockstep; asking for 33 costs the same as 64, and an if can halve throughput. Warp divergence serialises paths (two paths = half throughput); fix by reordering data so a warp wants one branch.
  • Occupancy hides memory latency, it is not a goal: a fast kernel at 25% beats a slow one at 75%. Raise it only when the SM stalls with nothing to run.
  • The memory ladder, each rung about 10x slower: registers ~1 cycle, shared/L1 ~30, L2 ~200, HBM ~450, host RAM over PCIe ~10,000. The whole game: move data up a rung and reuse it there.
  • Roofline: the ridge point is peak/bandwidth ~156 FLOP per byte; below it you are memory-bound and a faster maths routine changes nothing, only moving fewer bytes helps.
  • “GPU 100% utilised” means a kernel was resident, not that the maths was busy; a memory-bound kernel reports 100% while the arithmetic idles. Measure arithmetic intensity and achieved bandwidth.
  • Coalescing: HBM moves whole sectors; consecutive threads landing far apart force many transactions and can run 8x slower. Tiling: load a tile into shared memory once and let many threads reuse it.
  • Tensor cores only multiply small matrices (312 vs 19.5 TFLOP/s); dimensions not multiples of 8 or 16 force padding, so round a hidden dimension up to 64 or 128.
  • Operator fusion: four ops usually means four kernels each round-tripping through HBM; fuse the chain. FlashAttention is a tiled, fused attention with an online softmax that removes the N-by-N OOM without changing the arithmetic.
  • Getting a kernel: profile first, then torch.compile, then Triton for the one that matters; hand-written CUDA is a last resort. A 5% speedup instead of 40% means graph breaks (a leftover .item(), a print, a branch on tensor contents).

Training lab

watch a network learn, live

  • Training is three moves repeated: guess, measure wrongness as one number (the loss), nudge every weight downhill; identical at two dials or a trillion.
  • The learning rate is the step size and the most important number: too big diverges and the loss explodes. Loss to NaN means the rate is too high; divide it by ten. Use a schedule: big steps early, decayed late.
  • A neuron is a weighted sum squashed to a soft yes/no (one edge); depth is votes about votes, turning straight lines into curves and spirals.
  • Overfitting: 99% training, 62% live means it memorised the set; fix with more data, a smaller/regularised model, and early stopping. Both errors bad and more epochs doing nothing = underfitting, which needs capacity.
  • The U-curve: training error only falls, test error falls then turns up. Stop at the validation minimum and keep that checkpoint, not the last one.
  • Three splits: train to fit, validation to choose, test touched once; tuning until the test looks great leaks it and its score stops being a forecast.
  • Loss is continuous, accuracy is a threshold; the model becomes less wrong about its probabilities before any argmax flips. Trust the loss for progress, the metric for value.
  • Batching is a hardware decision: a batch turns thousands of small sums into one matrix multiply, reusing weights fetched once. Batch 8 to 1024 dropping quality means the averaged gradient lost useful noise; rescale the rate.
  • Fine-tuning ladder: prompt (change the brief), RAG (facts stay outside), fine-tune/PEFT (weights change). Facts go to retrieval; skills, style and jargon go to weights. A vocabulary gap is a weight gap.
  • LoRA on a 7B model: full fine-tuning wants over 100 GB (roughly 8x the model); freeze the base (14 GB at 16-bit, under 4 GB with QLoRA) and train adapters under 1% of parameters, fitting one 24 GB card.
  • Catastrophic forgetting: specialising on a small pocket redraws the whole map, so the damage is global. Freezing layers alone does not save you; the fix is in the data: mix general data back (replay), lower the rate, gate on a general test suite.
  • Distillation bottles a big teacher into a smaller student: training on a handful of hard labels just guesses; distilling on many teacher-labelled points with soft labels carries the boundary across, but the student must still pass the held-out test.

Before the network

four little machines

  • First ML question is never “which algorithm”; it is what do you have: labels (supervised), structure (unsupervised), or memories (lazy). Only supervised needs a gradient.
  • KNN is lazy: no training; every prediction measures distance to every stored example, so it pays at question time. Millions of rows need an approximate-neighbour index, the reason vector databases exist.
  • k=1 lets one mislabelled record own a region forever; use an odd k (5 to 15) so votes cannot tie; enormous k blurs into the majority. Tune k on held-out data.
  • Distance models treat every column’s units as comparable, so the widest-ranged column silently owns the metric; scale features before KNN and k-means.
  • Decision trees are unstable: one different split near the root cascades everywhere below; a forest of voting trees cures it, at the cost of the readability you built it for. Trees are greedy and cannot fix a bad early question.
  • A deep tree carves one rectangle per row and memorises noise; depth is capacity, so cap or prune and judge only on unseen data. But a tree is the one readable machine, worth several points of accuracy when a decision must be explained.
  • k-means settles in the nearest valley, not the best: each assign-then-update beat only shrinks total walking distance. Two seeds in one crowd split it forever (a trap ~78% worse than the ~1,268 optimum). Restart several times, keep the best.
  • Choosing k is your job: more clusters always cut distance, so pick where one more stops buying much. Unsupervised output is a proposal a human must name or reject.
  • Perceptron (1958) converges only when a perfect straight line exists; overlapping classes make it lurch forever and want the probabilistic line (logistic).
  • Of many perfect lines pick the widest margin: correctness is a pass mark, margin is the safety rating. PCA keeps spread, not meaning, so a low-spread direction flattens groups into one blob; try more than one angle.

LLM internals

one token at a time

  • A model does one thing: given tokens, output a probability for every possible next token; chat is that step looped with its output fed back. It does not plan or hold intent.
  • Which token is picked is chosen outside the model by a sampler you control; most “too random” complaints are sampler settings.
  • BPE tokenisation splits text into a fixed vocabulary (~50k to 130k) by merging the commonest byte pair; you are billed per token, about 4 chars each for English, nearer 3 for code, near 1 for Japanese. Cost tripling in another language is tokenisation, not degradation.
  • Attention is quadratic in context and the KV cache grows linearly, so long conversations getting slower is the expected shape, not degradation; fix with trimming, prefix caching or a shorter window.
  • The attention sink: many heads park surplus weight on the first token or two; dropping them destabilises softmax across the stack. Keep the first few tokens, trim from just after.
  • LayerNorm is the volume knob before every station: a 13% average per-layer gain over forty layers compounds to about 142x and overflows, so renormalising to a standard volume stops depth compounding.
  • Temperature controls variability, not correctness: at 0 you get the single most likely continuation every time, making a systematic error perfectly reproducible rather than absent.
  • Top-p (nucleus) beats top-k because the kept set adapts to confidence: close it once cumulative probability crosses p, then renormalise. A generous top-k keeps the junk tail so it wins some draws; top-p 0.90 shuts it out by the rule itself.
  • KV cache stores keys and values (they depend only on token and position), turning decode from quadratic into linear; without it a 4,000-token conversation costs forty times a 100-token one per generated token.
  • The cache trades arithmetic for memory and can exceed the weights; paging, eviction, prefix sharing, quantised caches and GQA all manage it, and paged attention removes the adjacency requirement via a block table.
  • Tool calling: the model emits a structured call (name plus arguments) instead of prose and executes nothing; your code validates and runs it, then feeds the result back. Function calling and tool calling are the same mechanism. Structure is not truth: a well-formed call can still be hostile, so validate every one.
  • Parameter census: at 124M the embedding tables are ~31% of the model; at 175B the same tables are ~0.4% and two-thirds is MLP. Weights dwarf activations, which is why serving one more user is nearly free.

LLM serving

the engine, not the model

  • Prefill and decode are opposite workloads: prefill reads the whole prompt in one compute-bound pass; decode makes one token per step, each reading the entire model, so decode is memory-bound.
  • Scheduling is iteration-level: the engine decides at every decode step what to admit, continue or evict; static batching waits for its longest request while finished slots idle.
  • Mean inside SLO but users say slow: look at the tail and split the metric. Queueing delay is non-linear, so P99 breaks long before the mean, and “slow” usually means TTFT.
  • Two latencies: TTFT (time to first token) and TPOT (time per output token); end-to-end = TTFT + TPOT x tokens, and the two move in opposite directions when you tune.
  • Bigger batch raises throughput and per-token latency together: it amortises the weight read but lengthens every decode step. The two goals genuinely conflict, so pick one and write it down.
  • Prefix caching: identical leading tokens share KV blocks (the system prompt is usually enormous); a collapsed hit rate means something variable is at the front (a timestamp, id, user name), so put stable text first.
  • Speculative decoding guesses several tokens then verifies in one pass: it helps a lightly loaded service and can reduce throughput under load; below roughly 40% acceptance the arithmetic does not work.
  • Multi-LoRA: keep adapters separate and apply at runtime so fifty tenants share one base and occupy the same batch; merging is right only for a single high-volume tenant.
  • Constrained decoding: asking for JSON gives JSON almost every time, but the failures are a trailing comma or a hallucinated field; guaranteeing JSON masks illegal tokens each step so they cannot be sampled.
  • Past ~90% utilisation, queueing delay grows faster than load, so the tail deteriorates sharply for a small efficiency gain: headroom is what you sell. Track goodput (requests per second that met their target), which cannot be gamed.

Fleet lab

twenty engines, one endpoint

  • An inference engine only knows its own GPU; it has no idea another replica holds the first two thousand tokens of your request.
  • Added replicas collapsing cache hit rate: the cache lives inside one replica and round-robin re-prefills a shared-prefix request somewhere that never held it. Route on prefix locality as well as load, which needs workers publishing what they hold.
  • Routing purely on cache locality creates a hot spot: every request sharing a system prompt piles onto one worker. It is a weighting between locality and load, not a mode; both extremes are misconfigurations.
  • Move the KV cache or recompute the prefix is arithmetic: transfer costs bytes over the interconnect, recompute costs prefill time. Intra-rack RDMA usually favours moving; ordinary Ethernet usually favours recompute.
  • Any worker list is stale the moment you write it: workers get rescheduled, restart after OOM, come up slowly loading weights. Discovery is eventually consistent.
  • Registered is not ready: a worker loading tens of GB of weights is running and useless, so it must not advertise as routable until it can serve. Separate liveness from readiness.
  • Planner oscillation comes from reacting faster than the feedback arrives: new capacity takes minutes to load, so corrections land late. Reduce the gain, add a cooldown, scale on queue depth or TTFT.
  • GPU utilisation is a bad scaling signal: it reads high whenever a kernel is resident, saturates early, is noisy, and does not fall when the queue drains.
  • Inference is not stateless: a mid-generation request holds a KV cache in one GPU, so when that worker dies the conversation dies. Retry safety depends on whether tokens have been sent: before the first, retry silently; after 400, an explicit error is the honest choice.
  • Capacity acceptance: sizing for peak is sizing for the day nothing goes wrong, and draining a node for a deploy is arithmetically the same as losing one. Size so the fleet still meets the target with one machine gone.