Build vs Buy: The Decision Framework Every Engineering Leader Avoids

You spent six months building the feature. The vendor launched it. Theirs is better.

This is not a rare outcome. It is the predictable result of a decision made without a framework. Build vs buy is one of the most consequential recurring decisions in any technology organisation. It is also one of the most consistently mis-analysed. The obvious costs — licensing fees for buy, engineering hours for build — are the easy costs. The hard costs stay hidden until it is too late to choose differently.

In 2010, Netflix moved its entire infrastructure to AWS instead of running its own data centres. The decision was controversial internally. The control argument was strong. Owning the hardware meant owning the reliability. But the engineering team ran the numbers. The cost of operating at their projected scale, including the talent required to run it, exceeded AWS by a margin that was not close. More importantly, operating data centres was not a differentiating capability. It was undifferentiated infrastructure. Building it was spending engineering capital on something that was not the reason Netflix existed.


The three axes

The decision has three axes: differentiation, total cost of ownership over the expected lifecycle, and risk. Each axis has a clear directional signal. Decisions that go wrong usually go wrong because one axis was ignored.

The primary frame is AT7 — Automation vs Control. Buying a capability accepts vendor control over its behaviour, roadmap, pricing, and availability. The control argument for building is real. Teams that build own the failure mode, the fix, the timeline, and the evolution. The automation argument for buying is also real. The vendor operates the capability at a scale most companies cannot replicate. The engineering cost is distributed across the vendor's customer base.

Both directions expose failure modes. A vendor dependency exposes FM1 — Single Point of Failure when the vendor is the only path to a capability. It exposes FM8 — Contract Violation when the vendor changes an API, deprecates a feature, or alters pricing without adequate migration time. Building exposes FM3 — Unbounded Resource Consumption. This is the ongoing cost of maintaining a capability as the underlying platform and security environment evolve.

The question is not "build or buy?" The question is: what is the cost of each at our scale, over our expected lifecycle, given our team's capabilities?


The decision matrix

Two questions, four quadrants.

Is this capability a core differentiator? Does a vendor solution exist that is good enough?

Buy when the capability is not core and a vendor exists. Email, auth, monitoring, log aggregation, payment processing. The vendor has more scale, more reliability investment, and more operational experience than your team can afford to build.

Build when the capability is core and no adequate vendor exists. Own the full lifecycle. This is your competitive edge.

Wrap when the capability is core but a vendor solution exists. Buy with an abstraction layer. Keep a migration path. This quadrant is the most dangerous because the correct answer requires discipline most teams lack. They either buy without the abstraction and lock themselves in, or they build from scratch and waste engineering capital on solved problems.

Partner when no vendor exists and the capability is not core. The market is immature. Co-develop or defer.

Buying core logic surrenders differentiation. Building undifferentiated infrastructure wastes engineering capacity. The matrix forces the questions build-vs-buy decisions most often skip.


The four forces distorting the decision

The force driving premature build is the engineer's default preference for building. Engineers find building interesting. Buying requires integrating someone else's design decisions, tolerating their limitations, and depending on their support quality. There is an aesthetic preference for building that is not always aligned with the organisational interest.

The force driving premature buy is the executive's default preference for speed. Buying feels faster. A contract is signed, the vendor delivers, the capability appears. The integration cost, the ongoing operational overhead, and the eventual migration cost are all deferred to a future that feels abstract in the procurement meeting.

The third force is lock-in risk. It is asymmetric. Teams overestimate lock-in for well-designed vendor solutions with standard interfaces. They underestimate it for vendor solutions with proprietary data formats or network effects.

The fourth force, rarely named, is political friction. Vendor relationships involve procurement, legal, and executive relationships independent of engineering merit. A technically inferior vendor may win because the vendor has a relationship with the CFO, or because procurement has already issued an RFP, or because a build decision would require headcount outside the current budget cycle. Technical leaders who do not name this force explicitly will find that decisions they believe are technical evaluations have already been made in a different room.


The costs that stay hidden

Hidden build costs: ongoing maintenance as the dependency landscape changes, security patching at a cadence the team did not plan for, hiring requirements to maintain a permanent team obligation, and opportunity cost. The engineering investment in building and maintaining the capability is not available for the capabilities that would differentiate.

Hidden buy costs: integration engineering, which is often larger than expected for solutions with complex APIs, migration cost when the vendor's solution no longer fits, feature roadmap risk since the vendor builds what benefits their largest customers, and dependency on the vendor's operational decisions and business continuity.

The team that answers both questions in financial terms before the architectural discussion begins is making a decision. The team that answers only the technical questions is making a recommendation that will be repriced later by someone else.


The reversibility test

Good technical leaders apply a reversibility test. If the build decision is wrong, how expensive is it to buy in two years? If the buy decision is wrong, how expensive is it to build or migrate in two years?

Decisions with low switching cost should be made faster. Decisions with high switching cost should be made more carefully, with the switching cost explicitly modelled.

They also separate the decision from the implementation. The question "should we build or buy?" must be answered analytically before "who would build it and how?" is even considered. Engineers who evaluate whether to build are simultaneously designing the solution. This biases the evaluation toward building.


The signal

When an engineering team spends more than twenty percent of its capacity maintaining a non-differentiating capability, evaluate the buy option. The maintenance cost is the integration cost amortised. You have already paid for the migration; you just have not called it that.

The most damaging failure is building what should have been bought during a growth phase, when engineering capacity was the scarce resource. Six months spent building email infrastructure produces a reliable email system and six months of lost product velocity. The infrastructure is maintainable. The products that were not built during those six months are absent from the market.

Now ask a harder question. Which of your current vendor relationships already resembles AT7 imbalance — you accepted control in exchange for operational simplicity, and the control cost is about to become real? What is your migration plan?

The full framework treatment — compression blocks, three-level exercises, and the complete AT/FM mapping — is in Book 6: The Architect's Mind, Chapter 9 (Build vs Buy). Free chapter available at computingseries.com/books/book6.