in

How Small Teams Should Staff, Scope and Run a First AI Build

The invoice that finally forces the decision is rarely the biggest one. It's the fourth or fifth SaaS renewal in a quarter, the one where a founder realizes the stack now costs more than a mid-level engineer and still can't answer the one question the business actually needs answered.

People frame the SaaS-versus-custom debate as a buying question, but the real work starts after a small team commits to building. Then come the harder decisions: who staffs it, how narrow to scope it, and how to run it without lighting cash on fire.

The failure math is unforgiving. MIT's NANDA research found that roughly 95% of enterprise generative AI pilots never deliver measurable financial returns. A small team doesn't get to be average on that curve. It has to run the build differently from the enterprise pilots that dominate the failure statistics, and the difference shows up in staffing, scope, and cadence long before it shows up in the model.

The SaaS Instinct Versus the Build Instinct

The SaaS instinct solves a new problem by adding another seat license. Fast, expensable, no hiring required.

The build instinct looks at the same problem and asks what it would cost to own the workflow outright. Both are legitimate. The trap is running on the SaaS instinct past the point where it stops being cheaper.

There's a breakdown of the subscription-stacking problem that walks through the crossover math in detail, and the pattern it names is worth internalizing: per-seat pricing scales with headcount, integration debt scales with the number of tools, and neither line item shows up on any single invoice. A custom build converts those variable costs into a fixed one. That trade only makes sense when the workflow is genuinely yours, and not when you're rebuilding a commodity.

Staff Small, Not Sparse

The temptation with a first AI build is to hire one senior engineer and call it a team. The temptation on the other side is to spin up a working group of eight because the project feels important. Neither works.

The research on team size has been consistent for decades — the Scrum Guide recommends 10 people or fewer, and most teams find their sweet spot between five and nine. For a first custom AI build inside a small business, the honest number is usually three to five, spanning a narrow set of roles rather than one heroic generalist.

Scope the First Version Ruthlessly

The big-bang AI project is the one that dies. The incremental approach — narrow, boring, measurable — outperforms the ambitious pitch almost every time, and it's the only version a small team can actually finance.

Pick one workflow. Pick the one where a wrong answer is recoverable and a right answer saves an hour a day per person. Ship that, then decide what's next.

Scope discipline lives in a document, not in someone's head. Write down what the first version does, what it explicitly does not do, and what would have to be true for a v2.

The document isn't for the engineers. It's for the founder who, four weeks in, will want to add "and also, could it draft the follow-up email?" Scope creep is the unspoken reason budgets double.

Run It Like a Product, Not a Pilot

The pilot mindset is what turns a promising build into another entry on the list of experiments that quietly died. Pilots have no owner after launch, no evaluation set, no plan for what happens when the model's answers drift.

Products do. The distinction sounds semantic and turns out to be operational.

Run the build in short cycles with something usable at the end of each one. Two-week iterations, a working artifact every time, and a real user testing it — not a slide showing what the model could theoretically do. Use hosted model APIs early so you're testing the workflow, not benchmarking infrastructure you may never need. If the workflow doesn't earn its keep at the API stage, no amount of custom training will save it.

And build the evaluation set before you build the model. A hundred real inputs with known-good outputs, curated by the person who does the job today, is the single most valuable artifact on the project. It's what lets you tell a good week from a bad one, and it's what keeps the team honest when a demo looks impressive and the numbers don't back it up.

Know When to Kill It

A small team's advantage over an enterprise pilot is the ability to stop. Set a kill criterion at the start: if the first version doesn't clear a specific bar by a specific date, it goes back on the shelf and the SaaS renewal stands for another year.

Naming that threshold up front is what separates a disciplined build from a slow-motion write-off. The teams that ship their first custom AI tool aren't the ones with the biggest budget. They're the ones willing to define, in writing, what failure would look like — and to act on it if they see it.

From a Better Lacrosse Ball to a Youth Sports Platform