Three ways to add a language model
The model is the engine under an AI feature: the assistant that answers from your documents, the agent that files a ticket, the classifier that sorts incoming email. The software around it (sign-in, screens, the search index, the tools the agent may use, the tests) is much the same whichever model you pick. What changes is how you pay for the model, who runs it, and what you can change about it.
| Route | What it is | What you pay for | What you control |
|---|---|---|---|
| Frontier | A leading commercial model, called through the provider's API. | Tokens: every word in and out, every month. | Your prompts, tools and data. Not the model, its versions or its retirement dates. |
| Open-source | A model whose weights are published, run on servers you choose. | GPU servers, every month, whether busy or quiet. | The model version, where it runs and where the data goes. |
| Custom | A model we make for you: fine-tuned, further trained, or trained from scratch. | Data work and GPU training time up front, then GPU serving. | Everything, including what the model knows and how it behaves. |
Most projects should start on the first route and move only when there is a reason to. The reasons are real, though, and they are worth knowing before you build.
Frontier models through an API
A frontier model is the most capable general model a provider sells. You send text (and often images or files) to an API and pay per token. There is no server to run, the provider improves the model over time, and you can have a working prototype in days.
The question most clients ask first is what happens to their data. The answer depends on the provider and the plan, so read the terms for the exact product you use. OpenAI states that data sent to its API is not used to train its models unless you opt in, and that abuse monitoring logs are kept for up to 30 days [1]. Anthropic states that by default it does not use inputs or outputs from its commercial products, including its API, to train its models [2]. Google's Gemini API terms draw a line between plans: content sent to the unpaid services may be used to improve Google's products, while prompts and responses on the paid services are not [3].
Frontier models also read a lot at once. Google lists an input limit of 1,048,576 tokens for Gemini 2.5 Pro [4], enough for several long documents in one request. A large context window is useful, but it is not a substitute for search: sending everything every time is slow and you pay for every token of it.
- Right when: you want the strongest reasoning available, your traffic is modest or uneven, and the provider's data terms suit your data.
- Watch for: token spend that grows with usage, models being retired on the provider's schedule, and rate limits at busy times.
- Plan for: an evaluation set you can rerun when the provider releases a new version, so an upgrade is a test result rather than a leap of faith.
Open-source models you run yourself
Open-weight models are published for anyone to download and run. You serve them on GPUs you rent from a cloud provider, or on a GPU server you buy. The data never leaves servers you control, the model never changes unless you change it, and the monthly cost is the servers rather than the tokens.
Open does not mean unconditional. Read the licence of the exact model and version. Meta's Llama 3.1 licence, for example, requires a separate licence from Meta for a licensee whose products had more than 700 million monthly active users on the release date, and asks you to display "Built with Llama" [5]. Other open-weight models are released under the Apache License 2.0, which grants a royalty-free licence to use, modify and distribute the work [6]. Your advisers should read the licence that applies before you ship.
- Right when: your data must stay in a place you choose, your traffic is steady enough to keep a GPU busy, or you need a model that will not change under you.
- Watch for: the GPU bill does not fall when nobody is using the product, and the strongest open models need large GPUs.
- Plan for: someone to update, monitor and secure the serving stack. It is a server like any other, with a model on it.
A custom LLM
A custom model is one changed or created for your use. There are three tiers, and the difference in effort between them is large.
- 01
Fine-tuning
An existing open model is trained further on examples of the task: your support replies, your report format, your classification labels. It learns a style and a skill, not a library of facts. This is the tier most custom projects need.
- 02
Continued pre-training
An existing model reads a large body of your domain's text before fine-tuning, so it learns the vocabulary and the background of a specialised field. It needs far more text and GPU time.
- 03
Training from scratch
A new model, built from your data and public data. It is the most expensive route by a wide margin and rarely the right one for a business. It makes sense when the model itself is the product.
Whatever the tier, the work that decides quality is the data and the evaluation, not the training run. A custom model is only as good as the examples it learned from, and you only know how good it is if you test it against questions it has never seen. GPU time for training is billed up front at cost, outside the milestones, because it is paid for before the work that uses it.
The same assistant, three ways
The three worked examples below are the same product: a web app where staff ask questions of the company's documents, and an agent that can look things up and open a ticket. Only the model route changes. Each shows the range from our rate card today, the build and the monthly running cost, priced when you open the page.
Worked example, priced now
On a frontier model
A document assistant with tools, on a frontier model through its API, at a medium level of use. The running cost is mostly tokens.
- Build
- ≈ US$89,000 to US$136,000, delivered within 21 weeksCAD 126,700 to 193,700
Prices in your currency are estimates from today's Bank of Canada rate. All invoicing is in CAD or USD.
Running it
- Hosting
- ≈ US$558 a monthCAD 795 a month
- Support
- ≈ US$1,000 a monthCAD 1,425 a month
- Model running cost
- ≈ US$383 a monthCAD 545 a month
Worked example, priced now
On an open-source model
The same assistant on an open-weight model served on rented GPUs at DigitalOcean. The running cost is mostly the GPU server.
- Build
- ≈ US$96,100 to US$147,000, delivered within 22 weeksCAD 136,800 to 209,100
Prices in your currency are estimates from today's Bank of Canada rate. All invoicing is in CAD or USD.
Running it
- Hosting
- ≈ US$4,360 a monthCAD 6,210 a month
- Support
- ≈ US$1,000 a monthCAD 1,425 a month
- Model running cost
- ≈ US$4,190 a monthCAD 5,960 a month
Worked example, priced now
On a custom fine-tuned model
The same assistant on a model fine-tuned on the company's own examples, then served on rented GPUs. The build includes the data work, the training time and the evaluation.
- Build
- ≈ US$133,000 to US$202,000, delivered within 33 weeksCAD 188,800 to 288,000
Prices in your currency are estimates from today's Bank of Canada rate. All invoicing is in CAD or USD.
Running it
- Hosting
- ≈ US$4,360 a monthCAD 6,210 a month
- Support
- ≈ US$1,000 a monthCAD 1,425 a month
- Model running cost
- ≈ US$4,190 a monthCAD 5,960 a month
Compare the running costs as well as the builds. A frontier model's cost moves with use, so a quiet month is cheap and a busy one is not. A served model costs about the same every month. If the frontier example's monthly cost at your expected traffic is well above the serving cost, that is the point at which the open-source route starts to pay. Change the traffic in the estimator to find it.
What stays the same on every route
Whichever model you choose, most of the project is the software around it, and that part does not change. The documents still need to be split, indexed and kept in sync. The agent's tools still need permissions, so it can look up an order but not refund one without a person. Answers still need their sources shown, so staff can check them. And the whole thing still needs an evaluation run before every change, logs that show what the model was asked and what it did, and a way to switch it off.
That is good news for the decision. Because the surrounding work is shared, the route is a smaller choice than it first looks, and one you can change later if the app talks to the model through a thin layer of its own. Choose the route that fits your data and your traffic today, and keep the option to move.
How to decide
- 01
Start with the data
Where may it go? If your advisers or your clients' contracts say it must not leave your servers or your country, the open-source or custom route decides itself. If a provider's paid API terms are acceptable, keep going.
- 02
Build the evaluation first
Write down fifty to a few hundred real questions with good answers. Every route is measured against the same set, so you choose on evidence.
- 03
Prototype on a frontier model
It is the quickest way to learn whether the feature is useful at all. Many projects stop here, rightly.
- 04
Price the running cost at real traffic
Estimate a month of real use. Compare tokens against a GPU server. The crossover depends on your volume, not on anyone's rule of thumb.
- 05
Try an open model on the same evaluation
If it scores close enough, you gain control and a flat bill. If it does not, you know what the frontier model is worth to you.
- 06
Consider custom last
Fine-tune only when the open model is close but not there, and you have the examples to teach it. Train from scratch only when the model is the product.
Mistakes we see
- Choosing a model before writing the evaluation. Without a test set, every model looks good in a demo.
- Fine-tuning to teach facts. Facts change; a fine-tuned model has to be retrained to learn them. Retrieval updates the moment the document does.
- Ignoring the running cost. A build that is cheap to make can be expensive to run, and the reverse.
- Building on one provider's special features so deeply that moving is a rewrite. A thin layer between your app and the model keeps the route a decision you can revisit.
- Treating an open-weight licence as a free-for-all. Each has its own terms.
We build on all three routes and price each on its own, so we have no reason to push one. When a frontier model through its API is the cheapest good answer, which it often is, that is what we will recommend.