Using AI in System Design Without Overengineering It

A five-layer reference pipeline, with the two gates in it that look sturdier on a diagram than they are in production marked honestly — because a guardrail that oversells its own reliability is worse than no guardrail at all.

Subject
AI in System Design Without Overengineering It
Published
19 AUG 2026
Reading time
9 min
In this post · 3 sections
  1. 01Start with the system, not the model
  2. 02Two techniques that get sold as more solved than they are
  3. 03The shape most reference architectures assume, and often shouldn't
The same argument in 3:32. Our graphics, AI narration.

Most AI-in-system-design advice fails in one of two directions: either AI touches nothing, out of caution, or it touches everything, because someone built a six-call agentic pipeline to do what a single classifier call would have done. Both are placement mistakes, not capability mistakes — the model was rarely the bottleneck. Where it sits in the system was.

McKinsey’s 2025 State of AI survey found 62% of organizations at least experimenting with AI agents — 39% experimenting, 23% already scaling an agentic system somewhere in the enterprise. In any single business function, no more than 10% say they are scaling agents at all. A separate McKinsey figure narrows that further: among organizations already using AI, only 7% say it’s fully scaled company-wide, which is the more honest number for how much of this is actually running in a way an organization can rely on rather than demo.

McKinsey, “The State of AI,” 5 November 2025. mckinsey.com

Start with the system, not the model

Before any AI decision: what are the actual inputs and outputs, where are the decision points, and which parts of the logic are genuinely deterministic versus genuinely uncertain? AI belongs at the uncertain parts — classification, generation, retrieval over unstructured data — and does not belong on a deterministic workflow, a critical path with no fallback, or a high-stakes decision with no human validation step. Naming which category a given piece of logic falls into, before reaching for a model, is most of the “without overengineering” problem solved in one step.

A five-layer AI system pipeline — input, process, model, validate, decide — with input filtering and the confidence gate flagged as weaker than they lookPIPELINEFIVE LAYERS, HONESTLYTWO GATES WEAKER THAN THEY LOOKDWG Nº 11DETERMINISM FIRSTSTREAMING — THE SAME FIVE STAGES, OVERLAPPED IN TIMEINPUTPROCESSMODELVALIDATEDECIDEREQUEST / RESPONSE — ONE STAGE AT A TIME, WHICH IS WHAT THE PIPELINE BELOW ASSUMESINPUTfiltered, not safe!PROCESSroute + prepareMODELnon-deterministicVALIDATEconfidence gate!DECIDEact or escalateDeterministic where you canAI only where it earns itValidate, don’t assume! WHAT THE DASHED GATES ACTUALLY AREINPUTA filter, not a fix — prompt injection is an open problem, not a checkbox.VALIDATEMost model APIs do not expose real calibrated confidence. Treat it as a hint.
Both dashed boxes are still worth having — a weak filter beats no filter, and a rough confidence signal beats none. The mistake is presenting either one as a solved problem on an architecture diagram, when the honest claim is that they reduce a risk they cannot eliminate. The MODEL box carries two marks, not one — Claude or Codex, whichever a team already runs — because the placement argument holds regardless of which one sits there.

Two techniques that get sold as more solved than they are

Semantic caching — serving a cached answer when a new query is similar enough to one you have already answered — is a real technique with a real risk that write-ups routinely skip: staleness. A cached answer is only correct for as long as the underlying data it was computed against stays correct, and the similarity threshold that decides “close enough to reuse” needs active tuning, not a default. Treat it as a cost optimization with a monitoring requirement attached, not a free win.

Confidence thresholds are the more consequential overstatement. Most LLM APIs do not expose a genuinely calibrated measure of task correctness — a token-level log-probability is not the same claim as “the model is 85% sure this answer is right,” and treating it as one is how a validation gate ends up passing wrong answers with high confidence. The honest framing is that a confidence signal narrows the review queue; it does not replace review.

On the injection line specifically: “sanitize the input” is not a fix, because prompt injection is not a known-shape attack the way SQL injection is — there is no complete filter, only mitigations that shrink the attack surface. The correct mental model is containment and blast-radius limits (what can this model actually do if it is fooled), not prevention.

The shape most reference architectures assume, and often shouldn't

A discrete input-process-model-validate-decide pipeline implicitly assumes batch or request/response workloads, where each stage genuinely finishes before the next begins. Streaming and real-time systems — voice, live chat — do not decompose this cleanly: validation may need to happen on partial output while generation is still running. If your system is one of these, the five stages above are still the right questions to ask, but they overlap in time rather than running as a strict sequence, and the architecture needs to be designed for that from the start rather than retrofitted.

AGNIZAR
Production AI · System architecture · Fractional CTO

Agnizar builds custom software and AI that fits the systems you already run, then hands it over or keeps it running. Every job starts small: one bounded piece of work, one named result, one clear decision. Tell us what you’re building; a senior engineer replies within a business day, not a salesperson.

Tell us what you’re building