Skip to content
PrasannaBrabourame

Forward Deployed AI Specialist · Singapore

A demo has to work once. Production has to work on the worst day of the quarter.

I'm Prasanna Brabourame. I join teams who are putting AI into work that has strict rules and gets inspected: licensing, compliance, security, schools. I build the whole thing — the screens people use, the AI behind them, the way it is kept safe, and the parts that keep working long after everyone has gone home.

Currently at NCS AI Central 10+ years shipping GovTech · RegTech · DevSecOps Open to advisory conversations
An illustrated portrait of Prasanna Brabourame, hand behind his head, surrounded by drawn diagrams: a cloud feeding a rack of servers, a chain of pipeline steps, a circuit-board brain, a padlock on a shield, a code window, a checklist and a dashboard of charts.
the work, and the person doing it

Most AI projects don't fail at the model.

They fail in the gap between something that looked good in a meeting and something a staff member can stand behind when an inspector asks. That gap is where I work.

A demo is judged on its best answer. A real system is judged on its worst one. The broken reply at three in the morning that nobody saw. The service refusing to respond in the middle of an inspection. The answer that sounded certain and was quietly wrong. I build everything assuming that day is already in the diary.

One pipeline. Pick a way production tests it.

requestguardin · schemaorchestratorplan, route, retrymodelguardout · policyanswerauditdecision tracetoolsallow-listedvector storecited, freshmemoryscoped, TTLworksone path, one time, in front of an audience429, and a cold startunder real concurrencybackoff · queue · warm pathfenced JSON, truncated array,a stray control characterschema · salvage · deterministic fallbacka retrieved document thatgives the agent instructionsprovenance · tool allow-list· output policythe window fills witheverything, ranked by nothingbudget · rerank · summarise · evictyesterday's session leakinginto today's answerscoped state · TTL · replayablea name, an NRIC, a salaryon its way out of the regionredaction · field controls · residencythe bill arrives, and it isthe same call ten thousand timesroute to the small model · cache · cap“why did it decide that?”rule referenceevidence · who overrode it

One question, one AI, one answer, with people watching. Everything below this line is a problem the demo never had to survive.

Right in the hour of an inspection, the AI service starts refusing requests because too many arrived at once. At the same time the serverless parts — small programs that only wake up when needed — are slow on their first run. Three answers. Wait and retry, with a spread-out delay called backoff. A queue that keeps everything in order. And a fast lane kept warm for requests that cannot wait.

The AI is asked for a tidy list and sends back something half-finished, or wrapped in stray characters. A strict format check (a schema) catches most of it, a repair step rescues what it can, and a plain fixed rule handles the rest. One bad reply spoils one line of one report. It never reaches the work waiting behind it.

Someone hides instructions inside an uploaded document, and the AI follows them as if they came from you. That is called prompt injection. Three defences. Every piece of text carries a note saying where it came from, its provenance. The AI can only touch things on an approved list, an allow-list. And a guardrail blocks any action nobody asked for.

AI can only read so much at once. Fill that space with everything you have, in no particular order, and the answers quietly get worse the more you give it. The fix has four parts. A token budget, meaning a firm limit on how much text goes in. Sorting, so the most useful material goes first. Summarising older material as you go. And a deliberate decision about what gets dropped, rather than letting it fall off the end.

Yesterday's conversation bleeds into today's answer, and afterwards nobody can work out how either one happened. Each conversation is kept separate, everything it remembers has an expiry time (a TTL) that really does expire, and any conversation can be replayed to see exactly what the AI knew at the time.

A name, an NRIC, a salary figure ending up somewhere it should never go: into a log file, into a message to the AI, or onto a computer in another country. Personal details are blacked out before anything is sent, which is redaction. Each field has its own rules. And where data is allowed to live, its residency, is enforced at the door instead of assumed.

The bill turns up, and it is the same question asked ten thousand times. Send the easy work to a smaller, cheaper AI, keep a copy of answers that never change so you only pay once, and set a spending cap that stops things before the invoice does.

An inspector asks why one particular decision was made, four months ago. Every decision keeps a record — a decision trace — holding the rule it came from, the evidence behind it, what the AI suggested, and the name of the person who changed it if anyone did.

Nine ways of looking at the same system. Only the first one fits in a demo.

The AI is one part, not the whole thing

When a decision affects someone's safety or has legal weight, it comes straight from the written rules: expiry dates, banned ingredients, how serious a problem is. The AI reads, sums up and suggests. It never gets the last word on anything a person has to answer for.

Broken replies are a normal Tuesday

AI sends back half-finished lists and stray characters all the time. So I build repair steps, and plain fixed rules to fall back on. The worst that happens is one wrong line in one report — not everything grinding to a halt.

It has to pick itself back up

Long jobs get cut off. Services start refusing requests. So something is always watching for work that has stopped moving: it restarts what it can, flags what it cannot, and turns a silent failure into a clear message a person can act on.

I trust AI exactly as far as the plain rule behind it can throw it.

Selected work

Four AI systems, in jobs where a wrong answer costs real money.

First idea to finished system in each case. The problems came from the people who do the work every day. The design, the building and the safety side were mine.

I have left out client names, product names and numbers. What follows is the shape of each problem and how it got solved, which is the useful part anyway.

National certification authority · Singapore Deployed

Certification document review

Officers used to read every supporting document page by page. Now the system works out what each uploaded file is, pulls the details out of it, and checks those details against the published rules. Instead of reading everything, the officer gets a list of what looks wrong and agrees or disagrees with each item.

documentsuploadedclassifyextractall extracted?one transaction, fires onceverifycross-checkofficer decidespass / warn / failwatchdog sweeps up anything that stalledthe model nevergets the last word

What I built

The screens the officers use — uploading, reading documents, working through the queue, seeing the history, printing a report — and everything behind them. Work moves along automatically as each step finishes, using two Singapore Government AI services, plus the part that holds all the written rules.

The barrier

Some checks compare documents against each other, so they can only run once every file is done. Uploads finish in any order, so the system uses a lock to make sure the final check runs once and only once per application. No duplicates, and nothing half-checked.

Where the model stops

Expired certificates and banned items are caught by a plain rule, never guessed at by the AI. Very long ingredient and menu lists are checked in small batches, and if a batch is too slow it splits in half again and again until it reaches one line — so a big list slows down instead of failing outright.

React 18TypeScriptTailwindFirebase Cloud FunctionsFirestoreAISayPlatform AI · GeminiSecret ManagerApp Check
AML/CFT compliance · Singapore Delivered

Regulatory compliance review

Checking regulated companies against the national rules on money laundering, from the desk rather than on site. An officer opens a case and uploads the paperwork: policies, background-check forms, director IDs, screening records. The system comes back with a list of problems, each marked for how serious it is, and the officer accepts, edits or rejects each one. Every change is kept on record.

evidenceuploadedextractper schemaLLM judgementreads, proposesrule-book checkdeterministictwo tracks, same rulefinding+ evidence + clauseofficersigns offseverity is stamped from the rule book — never inferred

The build

The overall design, the way the information is organised, the screens the officers work in, and the engine behind it. Every rule is checked two ways at the same time: once by the AI making a judgement, and once by plain checks on the actual fields. The final answer is built from both.

Traceable to a clause

If an inspector is going to act on something, they need to see exactly which rule it came from. How serious a problem is comes straight from the rule book, never from the AI's opinion. Every item carries its evidence, the rule it relates to, and every change an officer made to it.

The rest is what any system like this needs. Nothing is allowed unless it has been explicitly permitted. Analysts, approvers and administrators each see only what their job requires. Only certain file types and sizes can be uploaded. Administrators can edit the rule book and the document checklist themselves. And the finished review prints to a PDF.

React 19Python 3.12Firebase FunctionsFirestoreAISayGemini · LLMaaSNorthflankDocker
DevSecOps · vulnerability remediation Delivered

Vulnerability remediation pipeline

Security weaknesses get reported by several different scanning tools, each in its own format. The system tidies them into one list, removes duplicates, scores how risky each one is, and then walks it all the way through: raise a ticket, look into it, get approval, write the fix, test it, merge it. The fix is only ever submitted if the tests pass.

GitLabAWSCIde-dup+ risk scoreanalyseapprovehuman, or low-risk policyfix branchtestspassmergerequestfailback to a humannothing is pusheduntil the tests are green

Architecture

Built so every outside service plugs in and out: where the code lives, where tickets are tracked, how people get notified, which AI is used. Swapping any of them in does not mean touching how the fixing actually works.

Letting a model near a repository

Letting AI write code that reaches a real codebase. It only ever works on a separate copy, runs that project's own tests, and submits the change only if everything passes. Anything else goes to a person. Approval is required first, and only the lowest-risk categories are ever approved automatically.

Reporting

The same weakness often gets reported by several tools at once, so those get matched up into one item. Each is scored for risk, and tracked to show whether things are improving. Reports line up against the standard industry list of the ten most common weaknesses. Management gets PDF and PowerPoint summaries. There is an average time-to-fix figure. And a running total of what the AI costs, so nobody gets a surprise.

Node.jsExpressTypeScriptPostgreSQLPrismaReact 19ClaudeGeminiOpenAIOllama
Early-childhood education · Singapore Delivered

Early-years learning insight

Teachers write down what they notice about each child, and those notes sit in spreadsheets doing nothing. This system reads the spreadsheets, matches them up with the official curriculum and the school's own records, and turns them into a short, readable picture of how each child is getting on.

trackingspreadsheetsname resolverfour strategies before it gives upBigQuerycurriculumagent sequenceretrievesummariseevaluateper-childreadiness insightan unmatched row is a child with no insighta mismatched one is worse

Three tiers

Three layers, and the cloud setup around them. The engine, with two-step sign-in, spreadsheet reading and name matching. The app the teachers use. And the AI layer, which looks things up, sums them up, then checks its own work before anything is shown.

Names

Teachers do not spell children's names the way the official records do. The system tries four different ways of matching a name before giving up, because a row it cannot match is a child who gets nothing — and a row matched to the wrong child is worse.

Hardening

The network is split in two so the sensitive part is not reachable from outside, there are limits on how often anyone can call it, uploads are locked down, and personal details are blacked out of the logs. This is children's data. There is also a test setup for improving quality and cost away from live use, and a full handover pack at the end.

Go · GinNext.jsPython · FastAPIVertex AI Agent EngineFirestoreBigQueryPub/SubCloud Run

How I work

Embedded with the team that owns the problem.

Forward deployed means what it sounds like. I work inside your limits, your rules about data and your idea of what finished looks like — rather than from a contract someone wrote before anybody had seen the actual problem.

experiencethe screen the person actually usesagents & orchestrationretrieval, tools, memory, routing between modelsevaluation & guardrailsoffline harness, deterministic backstops, red-teaminggovernance & compliancepolicy, access, audit trail, data residency, explainabilitydata & integrationone record, many systems that never agreed on itplatform & deliveryinfrastructure as code, CI/CD, observability, costone pair of handswhere aprojectusuallyhands off

Pick a layer — or tab through them.

Every layer, and the points where most projects get handed to someone else.

Your use cases come before my architecture

Every system above started with the people who do the job telling me what they actually need. The design followed the limits they described: the regulations, the way they already review things, and the bits nobody ever writes down. The technical design came after that, not before.

I build all of it, top to bottom

The screens, the AI behind them, how the information is stored, how it is kept safe, the servers, and the way new versions get released. Fewer handovers to argue about, and nobody sitting waiting on anybody else. For a first build in an unfamiliar area, that is usually the quickest honest way to get something real.

Ready for the day after launch

Retrying properly when something fails, watching for work that has got stuck, repairing bad replies, plain rules to fall back on, a record of who did what, and passwords kept out of the code. This is the boring half of the job. It is also the half that decides whether the thing is still running in six months.

Built to be handed over

Written documentation, ready-made setup files, a security and tidy-up review, and training material for the people who will actually use it. The reasoning stays with your team, so the next thing they build should not need me at all.

Track record

Ten years of shipping, in three phases.

Research and language AI before it was fashionable, then building platforms and products, then back to AI — this time with real consequences attached.

research & NLPplatform & productAI in production2016 — 2019Senior Programmer, R&DIntegra Software2019 — 2022Senior Product EngineerLogical Steps2022 — 2023Team Lead (R&D)2359 Media2023 — 2024Tech LeadNCS2024 — 2026Senior ConsultantNCS Gov+2026 — nowForward DeployedNCS AI Centralcustom NER, NMT,a patent filingconversational AIsearch at scalemonolith → microservices,national platformsagentic AI officerscan defendPuducherry, IndiaSingapore
Ten years, three phases, two countries, and the thread running through them.

2019: Puducherry to Singapore, for a search engine.

2026 – now

AI Implementation Strategist

NCS AI Central · Forward Deployed Specialist · Singapore

Working inside Singapore Government AI programmes, alongside GovTech, Google and the National AI Group. Building with Claude, Gemini, Vertex AI Agent Engine and the Model Context Protocol.

2024 – 2026

Senior Consultant · Tech Lead & Cloud Architect

NCS Group · Gov+ programme

Designing the cloud setup and leading the technical work on government systems. ShiftRing AI: phone and chat customer support that understands ordinary speech, replies in several languages and raises tickets on its own. Safe at Work, which lets employers check the permit status of their staff.

2023 – 2024

Tech Lead

NCS Group · Singapore

Led the rebuild of FWMOMCare, the national health monitoring system for migrant workers. It went from one large program to a set of smaller connected ones, and moved onto a different database. Also Exit Pass, which handled dormitory exit permissions and location limits. And MW Data Hub.

2022 – 2023

Engineering Team Lead (R&D) & Principal Engineer

2359 Media · Singapore

Delivered the NUS IASS platform, which connected students to paid short-term work that counted towards their internship credit, on web, iPhone and Android. Also a complete relief-staffing system for PAP Community Foundation, covering everything from posting a job through to paying people.

2019 – 2022

Senior Product Engineer

Logical Steps · Singapore

Technical lead on the Ola search products. That meant a search engine that understood ordinary questions, a part that worked out what someone actually meant, a chatbot for voice and text, and a tool for tuning how results were ranked. Also LNDDO, which judged whether small businesses were creditworthy from their online activity. And CardsPe, which got shop owners selling online, built from nothing.

2016 – 2019

Senior Programmer · Innovation & Research

Integra Software Services · Puducherry, India

Designing systems and frameworks for the publishing industry. Built iNLP, which picked out names and terms from text, worked out which language something was in and sent it to be translated, removed duplicates across huge volumes, and adapted existing AI models to new tasks. Designed iAuthor and iCorrectProof from scratch while also running the team, and built a page-layout engine that ran in the browser — the research behind it led to a patent application.

publishing & NLPiNLP, iAuthorfintechLNDDO, CardsPeGovTechFWMOMCare, Exit PassRegTechcertification, AML/CFTEdTechearly years, gig workDevSecOpsremediationdifferent industries, the same underlying problem:high-consequence decisions buried in documents
Six sectors, one recurring problem.

Education

B.Tech, Electronics & Communications Engineering
Pondicherry University, 2012–2016

Recent certification

Anthropic: Claude Platform, Claude Code, Model Context Protocol (incl. Advanced Topics), Agent Skills, AI Fluency. Microsoft Applied Skills for AI research agents. NUS-ISS ICT Assessment for Software Developer.

Also

Best Team, Global DevOps Bootcamp at Microsoft. Thirteen published npm packages. Writes on AI engineering and production war stories at Medium. English and Tamil.

Capabilities

What I bring to the first week.

ClaudeGeminiVertex AgentMCPRAGagenticAI & LLMTypeScriptPythonGoNodeC#languagesGCPFirebaseAWSCloud RunDockerK8scloudPostgresFirestoreBigQueryPub/SubdataReactNext.jsTailwindfront endRBACauditsecretsrate limitsredactionhardeningCI/CDTerraformhandoverdelivery

AI systems

Building AI that carries out multi-step jobs on its own, across more than one provider so you are not locked to any of them: Claude, Gemini, OpenAI, Bedrock and Ollama. Tools for that include Vertex AI Agent Engine, Google ADK, Model Context Protocol, LangChain and LangGraph, and CrewAI.

Most of that is now multi-agent orchestration: several narrow agents doing one job between them, rather than one agent trying to do everything. Above them sit guardian agents, whose only task is to check the others before a person sees the result. Agentic coding too, where the AI writes changes that reach a real codebase.

Also letting AI answer from your own documents, known as RAG, and sorting and reading uploaded files. There is always a plain rule behind the AI to catch it. Earlier on: spaCy, Hugging Face, TensorFlow, Rasa.

Context engineering

Deciding exactly what the AI sees on every call: the instructions, the documents pulled in, the history kept, the tools offered. Most agent failures are not the model being stupid. They are the model being handed the wrong things. It is the biggest lever on the bill too, because an agent can send hundreds of thousands of words in and get a paragraph back.

What it costs and how fast it feels are things you design, not things that just happen. Prompt caching, so you do not pay twice for the same question. Model routing, so easy work never reaches an expensive AI. Caveman prompting, which means cutting instructions down to the few words that actually change the answer. Deliberate context headroom, leaving spare room so replies never get cut off at the edge. And an LLM wiki: a tidied-up summary of your knowledge that the AI reads, instead of wading through every original document. Plus putting the most useful information first, shortening what gets sent, summarising as you go, grouping requests together, asking for answers in a fixed format, and using a smaller, cheaper AI wherever it is honestly good enough.

Knowing whether it works

If you cannot measure a change, you are only guessing that it helped. So: DeepEval, Ragas, Promptfoo, LangSmith, TruLens, Arize Phoenix and W&B Weave. Checking not just whether the AI reached a believable answer but whether it got there sensibly. Using one AI to mark another's work, with a plain rule to settle close calls. Keeping a set of known-correct examples, and a test harness that runs against them every time before a change goes live — plus a way to trade quality against cost away from real users.

Building

TypeScript, Python, Go, Node.js and C#. React and Next.js for the screens people see; Express, NestJS and FastAPI for the engine behind them. PostgreSQL, Firestore, MongoDB, BigQuery, Redis and Elasticsearch for storing and searching data. Systems where each step starts as soon as the last one finishes, servers that only run when needed, small services instead of one big program — and, now and then, splitting up a big old program that has outgrown itself.

Cloud & delivery

Google Cloud and Firebase, AWS and Azure. The usual building blocks: running code without managing servers, passing messages between parts, keeping passwords safely, running jobs on a schedule. Packaging everything so it runs the same everywhere, describing the whole setup in files rather than by hand, releasing new versions automatically, and working within the rules that apply when the data has to stay inside the country.

Doing it safely

Nothing is allowed unless it has been explicitly permitted. People see only what their role needs. Stored data is scrambled so it is useless if stolen. Every request has to prove who it is. Uploads are locked down and there are limits on how often anyone can call the system. Personal details are blacked out of the logs, and there is a full record of who did what. Plus the documents and training that let someone else take it over completely.

Certification & continuous learning

The tools change every quarter. Staying fluent is part of the job.

None of this replaces having actually delivered something. It is how I keep up with tools that did not exist a year and a half ago. It also happens to leave a record of what I was learning, and when.

13Feb5Mar5Jun15Jul2026search, data stores,enterprise assistantsagents, tools, memory,guardrails, evaluationthe count is not the point — the subject changedGoogle Cloud skill badges earned per month

Google Cloud

Diamond League · 27,897 points
Two certifications this month. Professional Cloud Architect on 26 September, which is the hard one, and Generative AI Leader on 9 September. Alongside them, 25 skill badges this year. Those cover the Agent Development Kit from end to end: building agents, putting them live, then measuring and improving them. Also running several agents together on one job. Gemini Enterprise, from a first application through to governing what its agents may reach. And the groundwork underneath: Vertex AI Search and its data, serverless applications on Cloud Run, handling data from the command line, and the cloud foundations track.
Public profile →

Anthropic

Claude Platform 101 · Claude Code 101 and Claude Code in Action · Building with the Claude API · Model Context Protocol, including Advanced Topics · Agent Skills · AI Fluency · Claude with Vertex AI · Claude Cowork.
Earned March – June 2026.

Verified badges

38 on Credly across Google Cloud, AWS and IBM, going back to 2018. Elsewhere: Microsoft Applied Skills for AI research agents, NUS-ISS ICT Assessment for Software Developer, AWS Cloud Practitioner Essentials, Machine Learning (Stanford), IBM Certified Data Architect – Big Data.
Verify on Credly →

Open resource

The notes are public.

Everything I revise from lives in one place, free and open to anyone: 267 topics across Kubernetes, Google Cloud, large language models, retrieval, agents, multi-tenancy, architecture and leadership. Each one is a card you can drill, sorted by how hard it is and by subject.

It also has a cloud lab that follows a single request from a phone all the way to the answer and back, across AWS, Google Cloud and Azure, and lets you break each layer to see what stops working.

Open the learning guide →

A short self-check

Is your AI ready to be depended on?

Five questions — the same ones I would ask in the first fifteen minutes of a call. There is no score at the end, just an honest look at where you stand.

Start with the awkward problem.

Fifteen minutes. Tell me about the thing that keeps not working: the pile of manual checking nobody can keep up with, the AI trial that never got past the demo, the system nobody trusts enough to act on. I will tell you honestly whether it is worth building and what I would try first. No slides, no proposal.