On-premise spec search assistant
Hypothetical. Not a client, not a result.
// The problem
Engineers search thousands of specifications, drawings and vendor documents to answer one question, and contract terms mean those documents cannot go to an outside model.
// What we would build
An assistant on an open-source model running on a GPU server in your own data centre that answers from specifications, drawings and vendor data with the page it came from, prepares MTO checks and flags inconsistencies between documents for an engineer to check.
// What is in it
- Specifications, datasheets and vendor documents indexed and kept in sync
- Answers that cite the document and page they came from
- MTO checks drafted against the specs, for an engineer to check
- Flags where two documents disagree on a value
- An evaluation set from real questions, run before every change
- An open-source model on a GPU server you own, with nothing sent outside
// Stack
- Llama or Qwen (open-source)
- vLLM
- PostgreSQL with pgvector
- Next.js
- Lenovo GPU server on premise
// estimate
- Build
- ≈ US$148,000 to US$225,000, delivered within 27 weeksCAD 210,600 to 320,600
- Hosting
- ≈ US$176 a monthCAD 250 a month
- Support
- ≈ US$2,500 a monthCAD 3,565 a month
- Hardware
- ≈ US$47,800 to US$197,000CAD 68,000 to 280,800
Prices in your currency are estimates from today's Bank of Canada rate. All invoicing is in CAD or USD.
- AI route
- Open-source model
- Model running cost
- ≈ US$176 a monthCAD 250 a month
Timeline by milestone
- Discovery
- 2.9 to 3.4 weeks
- Specification and evaluation plan
- 0.4 weeks
- Design approved
- 2.4 to 3.9 weeks
- Core features
- 1.4 to 2.4 weeks
- Full build
- 0.9 to 1.4 weeks
- Working pilot
- 5.9 to 10.4 weeks
- Testing and fixes
- 1.4 to 2.4 weeks
- Launch
- 0.9 to 1.4 weeks
- Production
- 0.9 to 1.4 weeks
Outside our hands, and added to the calendar
- Hardware delivery, after it is ordered
- 2 to 6 weeks
- Client IT security review and vendor onboarding
- 4 to 8 weeks
How it is paid
- Deposit 20%
- CAD 42,060 to 64,060
- Discovery 1.6%
- CAD 3,364.80 to 5,124.80
- Specification and evaluation plan 11%
- CAD 23,133 to 35,233
- Design approved 1.6%
- CAD 3,364.80 to 5,124.80
- Core features 4.9%
- CAD 10,304.70 to 15,694.70
- Full build 3.2%
- CAD 6,729.60 to 10,249.60
- Working pilot 33.1%
- CAD 69,609.30 to 106,019.30
- Testing and fixes 1.6%
- CAD 3,364.80 to 5,124.80
- Launch 1.6%
- CAD 3,364.80 to 5,124.80
- Production 11.4%
- CAD 23,974.20 to 36,514.20
- Holdback, 30 days after launch (10%)
- CAD 21,030 to 32,030
- GPU time for training and testing, at cost
- CAD 300. Billed up front, at cost, outside the milestones
- GPU server for AI, for the model
- CAD 68,000 to 280,800. Billed up front, at cost, outside the milestones