Blog
Engineering notes on production AI and software architecture.
What we’ve learned building production AI, system architecture, and the operating decisions that carry either one from a working idea to real users. Concrete over abstract, sourced over asserted — every number here is either real and linked, or it isn’t stated.
Local AI on a Mac: How Much Memory You Need. Own the Search; Rent Judgement at About 2 Cents per 1,000 Calls.
Memory decides how much local AI a Mac can run, from 8GB to 512GB. Google’s EmbeddingGemma 300M fits every Mac for search; on routing decisions TypeSafe’s Jev 1.13 scored higher, for about 2 cents per 1,000.
Jev 1.13 Tested: Its Router’s Accuracy for 2 Cents per 1,000 Decisions
TypeSafe’s Jev 1.13 only decides; it never writes. Picking AI model sizes for 94 of our tasks, it scored 73 to its own router’s 72, in a fifth of a second and for about 2 cents per 1,000 decisions.
Sovereign AI: Three Questions Every Business Must Answer
Enterprises are demanding sovereignty over their AI systems to protect proprietary data, ensure operational continuity, and comply with strict regulatory mandates. Here is why the movement is accelerating, and the three strategic questions leadership must answer.
A Second Reader Catches What the First One Missed. Not What Retrieval Missed.
Two models answering independently from the same retrieved passages catch each other’s misreadings, because their mistakes are not the same mistakes. What neither of them can catch is a passage that should never have been retrieved — and that is the outcome the mechanism reports as healthy.
The Cheapest Model That Clears Your Eval Bar Is the Correct Model
Most teams pick a frontier model in the kickoff meeting and never revisit it. That is not a technology decision — it is the absence of one. Here is what a routing harness actually looks like, and the three places a cascade quietly breaks.
Your Context Window Did Not Kill RAG. It Just Moved the Failure Mode.
Bigger context windows were supposed to make RAG unnecessary. The research says models read long inputs unevenly, length alone can cost accuracy, and routing between the two can keep quality at lower cost.
Agentic vs Generative AI: What Actually Belongs in Your System
The distinction is real and worth keeping straight, but the decision that matters more than either label is what happens the second time a tool call fires — because the first time is never the one that breaks production.
Product Development Isn’t a Straight Line, and Ignoring the Loops Costs You
Product development is often drawn as a straight line of stages, but validation and testing routinely send work backwards. Those loops are cheapest when they come early, and a stable product with proven demand rarely triggers them.
Using AI in System Design Without Overengineering It
A five-layer reference pipeline, with the two gates in it that look sturdier on a diagram than they are in production marked honestly — because a guardrail that oversells its own reliability is worse than no guardrail at all.
The Strangler Fig Pattern Replaces a Legacy System Without a Rewrite
The strangler fig pattern replaces a legacy system one piece at a time, with the old system staying live as the fallback. Tests, not a launch date, decide when each piece is safe to switch over.
Redis and Terraform Left Open Source the Same Way. Only One Came Back.
Redis and Terraform both left open source for a more restrictive licence within months of each other, and outside groups copied each one within weeks. Only Redis came back; for Terraform, the copy became the lasting open version.
Enterprise AI Strategy Is Five Decisions, Made in the Right Order
An enterprise AI strategy is five decisions, usually made out of order or skipped. Skip governance or monitoring, and AI goes live before anyone can show it’s safe to run.
The Five Gates an AI Tool Has to Clear Before You Buy It
Not a ranked list of tools — those go stale within months. Five questions that do not, including the one almost every roundup skips: what happens to your data, and what it actually costs once you are running five of these at once.
Conway’s Law: Your System Is Shaped by Who Actually Talks to Whom
Conway’s Law says a system’s architecture ends up mirroring how its teams communicate, whatever the org chart on the wall says. The fix is mechanical: make approvals, deploy access and the on-call rota match how teams work.
AI Integration Projects Need a Check on the Output, Not a Better Model
AI integration projects mostly stall on workflow, not on the model, according to MIT’s 2025 study. Before scaling one, build a check that catches a wrong answer before it reaches a customer, a regulator or print.
Your CMS Architecture Is a Revenue Decision, Not a Content One
A content-management system (CMS) reads like an editorial choice, but its architecture decides revenue: whether rankings survive a migration, whether the page appears without waiting for the ad auction, and whether you can switch back if launch goes wrong.
Your Voice Agent Doesn’t Have a Model Problem. It Has a Wiring Problem.
When a voice agent feels laggy, the instinct is to swap in a faster model. In a typical stitched pipeline, the model is rarely where the time goes — the network hops between separately-hosted speech, reasoning and synthesis services are. Here is the 800-millisecond budget, spent honestly.
Cognition and Anthropic Don’t Actually Disagree About Multi-Agent Systems
Two posts from two frontier labs, one day apart, read as opposite advice: don’t build multi-agent systems; multi-agent systems cut our research time by 90%. They’re both right — they’re describing different shapes of task. The difference is a decision rule, not a preference.
MVP, MLP, MMP: The Names Don’t Matter. What Changes Under the Hood Does.
Every stage in this ladder gets defined by what a user sees. The expensive part — what a system is allowed to fake, and what it has to stop faking — is architectural, and almost never written down.
Build, Buy, or Partner: How to Actually Invest in AI Capability
Most “how to invest in AI” advice is written for a stock portfolio, not an engineering budget. Here is the real decision — three paths, three different cost structures — and an honest look at how much of AI’s promised potential companies have actually captured.
The Half-Life of an Engineering Decision: Institutional Knowledge and Continuity
Every workaround has a reason nobody wrote down. When the person who knows it leaves, the reason leaves first — the code stays behind and quietly stops making sense. What continuity is actually worth, in mechanism rather than platitude, and the honest case for when it isn’t.
What a Fintech App Actually Costs Is the Reconciliation Logic, Not the Screens
Cost estimates for fintech products spend their time on features and compliance categories. The part that actually blows budgets — making sure two concurrent payment requests can never double-charge the same account — rarely gets a line item at all.