How we pick an AI model for a client project

761 VenturesOctober 6, 20264 min readDelray Beach, FL

No client has ever asked us which AI model is inside their assistant. They ask whether it answers right, how fast, and what it costs to run. The model choice follows from those three.

Short answer

We pick a model by the job it has to do, not by whatever is newest. Short answers from a website need a fast, inexpensive model. Reading long documents or writing careful drafts needs a deeper one. Most of what we ship uses more than one model, each doing the part it is suited for.

Start with the job, not the model

Before we look at any model we write down what the thing has to do in one sentence. "Answer questions from this website and collect a callback." "Read an applicant's answers and recommend hire or pass." "Turn a permit document into a checklist." "Write a first draft of a weekly post using facts from the owner's notes." Each of these is a different job, and the same model is rarely the right fit for all of them.

Then we sort the job on a few lines:

What we reach for, as of this writing

The model families we work with most are OpenAI's GPT, Anthropic's Claude, Google's Gemini, Meta's Llama, xAI's Grok, DeepSeek, and Mistral. Each has a reputation among builders, and the reputation shifts every few months, so take this as how we think rather than a scoreboard.

For careful writing, following long instructions, and reading documents, we lean toward Claude. For broad tool support and the widest set of ready made integrations, GPT is often the quick path. Gemini tends to fit when the product already lives in Google's world or needs to handle images and long inputs together. Llama, Mistral, and DeepSeek are the families we run ourselves when the data cannot leave a client's control or when the volume is high enough that owning the cost makes sense. Grok we use less often but keep an eye on.

None of this is a ranking. The same family can be the right call for one client and the wrong call for the business next door.

Why one product usually has more than one model

A website assistant we build is often three models working in a line. A small, fast one reads the visitor's message and decides what kind of question it is. A mid sized one writes the answer from the site content. A careful one runs once at night to summarize the day's conversations and draft the owner's recap. Using the careful model for every message would be slow and expensive. Using the fast one for the recap would produce a thin summary. Each does the part it fits.

The hiring screener works the same way. One model runs the live phone conversation, where speed matters. A different one reads the transcript afterward and writes the recommendation, where care matters more than speed.

How we test before we commit

We do not pick from a spec sheet. We take a stack of real inputs from the client (real visitor questions, real applicant answers, real documents) and run them through two or three candidates. Then we read the outputs side by side with the owner. The one that sounds like their business and gets the facts right wins, even if it is not the one with the most buzz that month.

What happens when the model changes

Models get replaced. Providers retire old ones and release new ones on their own schedule. We build so the model is a setting, not a foundation. Swapping it is a configuration change and a test run, not a rebuild. We keep the test inputs from the original selection so when a new option appears we can run the same stack and see if it does better.

This is the quiet part of what we build: the assistant on a landscaper's site, the screener for a builder's hiring, the brief an owner reads each morning. The model is a decision we make per job, and we are prepared to make it again.

Questions people ask

Do you use the same AI model for every client?

No. The model is chosen per job, and most client systems use more than one. A fast model handles live chat, a deeper one handles document reading or careful writing, and sometimes a self hosted one handles anything private.

What happens when a newer model comes out?

We build every product so the model is a setting rather than a foundation. Swapping it is a configuration change followed by a test run against the same real inputs we used when we first chose. If the new one does better, it goes in.

Does the client need to know which model is being used?

They can always ask and we always tell them, but it is rarely the question that matters. What matters is whether the answers are right, how fast they arrive, and what it costs to run. The model choice follows from those.

Share
Written by the 761 Ventures team

Operators in Delray Beach, Florida who build websites, CRMs, AI assistants and automation for businesses like yours, then write down what worked.

Want this done for your business?

Two minute intake. A real person reads every one and replies within a business day.

Start a project
OlderWhat a law firm website needs to get found in Palm Beach CountyNewerWhat an AI assistant on a website actually does all day

Keep reading