How we pick an AI model for a client project
No client has ever asked us which AI model is inside their assistant. They ask whether it answers right, how fast, and what it costs to run. The model choice follows from those three.
We pick a model by the job it has to do, not by whatever is newest. Short answers from a website need a fast, inexpensive model. Reading long documents or writing careful drafts needs a deeper one. Most of what we ship uses more than one model, each doing the part it is suited for.
Start with the job, not the model
Before we look at any model we write down what the thing has to do in one sentence. "Answer questions from this website and collect a callback." "Read an applicant's answers and recommend hire or pass." "Turn a permit document into a checklist." "Write a first draft of a weekly post using facts from the owner's notes." Each of these is a different job, and the same model is rarely the right fit for all of them.
Then we sort the job on a few lines:
- Speed versus depth. A website assistant has to respond while someone is still looking at the screen. A document reader can take its time in the background. Faster models are lighter; deeper models think longer.
- How much it reads at once. A short chat needs a little context. A permit packet or a long transcript needs a model that can hold a lot of text in view without losing the thread.
- Cost sensitivity. An assistant that answers many visitors a day runs up a bill in small pieces. A screener that runs a few times a week does not. The volume shapes what we can afford per call.
- Privacy. Applicant records, financial details, and anything a client would not want leaving their building get stricter handling, sometimes a different provider, sometimes a model we host ourselves.
- How bad a wrong answer is. A clumsy phrasing in a chat is fixable. A wrong hire or pass recommendation, or a made up fact in a published post, is not. The higher the stakes, the more we pay for care and the more review we add.
What we reach for, as of this writing
The model families we work with most are OpenAI's GPT, Anthropic's Claude, Google's Gemini, Meta's Llama, xAI's Grok, DeepSeek, and Mistral. Each has a reputation among builders, and the reputation shifts every few months, so take this as how we think rather than a scoreboard.
For careful writing, following long instructions, and reading documents, we lean toward Claude. For broad tool support and the widest set of ready made integrations, GPT is often the quick path. Gemini tends to fit when the product already lives in Google's world or needs to handle images and long inputs together. Llama, Mistral, and DeepSeek are the families we run ourselves when the data cannot leave a client's control or when the volume is high enough that owning the cost makes sense. Grok we use less often but keep an eye on.
None of this is a ranking. The same family can be the right call for one client and the wrong call for the business next door.
Why one product usually has more than one model
A website assistant we build is often three models working in a line. A small, fast one reads the visitor's message and decides what kind of question it is. A mid sized one writes the answer from the site content. A careful one runs once at night to summarize the day's conversations and draft the owner's recap. Using the careful model for every message would be slow and expensive. Using the fast one for the recap would produce a thin summary. Each does the part it fits.
The hiring screener works the same way. One model runs the live phone conversation, where speed matters. A different one reads the transcript afterward and writes the recommendation, where care matters more than speed.
How we test before we commit
We do not pick from a spec sheet. We take a stack of real inputs from the client (real visitor questions, real applicant answers, real documents) and run them through two or three candidates. Then we read the outputs side by side with the owner. The one that sounds like their business and gets the facts right wins, even if it is not the one with the most buzz that month.
What happens when the model changes
Models get replaced. Providers retire old ones and release new ones on their own schedule. We build so the model is a setting, not a foundation. Swapping it is a configuration change and a test run, not a rebuild. We keep the test inputs from the original selection so when a new option appears we can run the same stack and see if it does better.
This is the quiet part of what we build: the assistant on a landscaper's site, the screener for a builder's hiring, the brief an owner reads each morning. The model is a decision we make per job, and we are prepared to make it again.
Questions people ask
Do you use the same AI model for every client?
No. The model is chosen per job, and most client systems use more than one. A fast model handles live chat, a deeper one handles document reading or careful writing, and sometimes a self hosted one handles anything private.
What happens when a newer model comes out?
We build every product so the model is a setting rather than a foundation. Swapping it is a configuration change followed by a test run against the same real inputs we used when we first chose. If the new one does better, it goes in.
Does the client need to know which model is being used?
They can always ask and we always tell them, but it is rarely the question that matters. What matters is whether the answers are right, how fast they arrive, and what it costs to run. The model choice follows from those.
Want this done for your business?
Two minute intake. A real person reads every one and replies within a business day.
Keep reading
What an AI assistant on a website actually does all day
What a website AI assistant does all day for a Florida service business: page matched openers, tap qualifying, price callbacks, CRM leads and a morning recap.
AI at workOpenAI, Anthropic, Google and the rest: what each is for when you are building
A builder's map of OpenAI, Anthropic, Google, Meta, xAI, DeepSeek and Mistral: what each is known for, hosted versus open models, and how we choose per job.
AI at workAgents versus assistants: the difference that matters to a business owner
Assistants answer, agents act.