Using AI in System Design Without Overengineering It
A five-layer reference pipeline, with the two gates in it that look sturdier on a diagram than they are in production marked honestly — because a guardrail that oversells its own reliability is worse than no guardrail at all.
- Subject
- AI in System Design Without Overengineering It
- Published
- 19 AUG 2026
- Reading time
- 9 min
- Film
- Watch · 3:32
In this post · 3 sections
Most AI-in-system-design advice fails in one of two directions: either AI touches nothing, out of caution, or it touches everything, because someone built a six-call agentic pipeline to do what a single classifier call would have done. Both are placement mistakes, not capability mistakes — the model was rarely the bottleneck. Where it sits in the system was.
McKinsey’s 2025 State of AI survey found 62% of organizations at least experimenting with AI agents — 39% experimenting, 23% already scaling an agentic system somewhere in the enterprise. In any single business function, no more than 10% say they are scaling agents at all. A separate McKinsey figure narrows that further: among organizations already using AI, only 7% say it’s fully scaled company-wide, which is the more honest number for how much of this is actually running in a way an organization can rely on rather than demo.
McKinsey, “The State of AI,” 5 November 2025. mckinsey.com
Start with the system, not the model
Before any AI decision: what are the actual inputs and outputs, where are the decision points, and which parts of the logic are genuinely deterministic versus genuinely uncertain? AI belongs at the uncertain parts — classification, generation, retrieval over unstructured data — and does not belong on a deterministic workflow, a critical path with no fallback, or a high-stakes decision with no human validation step. Naming which category a given piece of logic falls into, before reaching for a model, is most of the “without overengineering” problem solved in one step.
Two techniques that get sold as more solved than they are
Semantic caching — serving a cached answer when a new query is similar enough to one you have already answered — is a real technique with a real risk that write-ups routinely skip: staleness. A cached answer is only correct for as long as the underlying data it was computed against stays correct, and the similarity threshold that decides “close enough to reuse” needs active tuning, not a default. Treat it as a cost optimization with a monitoring requirement attached, not a free win.
Confidence thresholds are the more consequential overstatement. Most LLM APIs do not expose a genuinely calibrated measure of task correctness — a token-level log-probability is not the same claim as “the model is 85% sure this answer is right,” and treating it as one is how a validation gate ends up passing wrong answers with high confidence. The honest framing is that a confidence signal narrows the review queue; it does not replace review.
The shape most reference architectures assume, and often shouldn't
A discrete input-process-model-validate-decide pipeline implicitly assumes batch or request/response workloads, where each stage genuinely finishes before the next begins. Streaming and real-time systems — voice, live chat — do not decompose this cleanly: validation may need to happen on partial output while generation is still running. If your system is one of these, the five stages above are still the right questions to ask, but they overlap in time rather than running as a strict sequence, and the architecture needs to be designed for that from the start rather than retrofitted.
Agnizar builds custom software and AI that fits the systems you already run, then hands it over or keeps it running. Every job starts small: one bounded piece of work, one named result, one clear decision. Tell us what you’re building; a senior engineer replies within a business day, not a salesperson.