Skip to content
AtheronLABS

You're visiting from the United States. Prices are shown in US dollars. Not right?

labs@atheron:~/industries/media/ai$ train agent --grounded

Archive search AI

The right clip on deadline, with its rights shown.

An archive is worth what a producer can find in it on deadline. We transcribe and index it and put an assistant over it that links straight to the clip, on open-source models you control, because unreleased and rights-restricted material should not leave your hands.

// what matters here

What is different in this industry.

01

Down to the moment

Search finds the timecode where a subject was discussed, not just the programme it was in.

02

Rights beside every result

Each clip shows its rights notes, so a producer knows what can air before they build a segment around it.

03

Private by design

Open-source models run on GPUs you control. Unreleased programmes and restricted footage are never sent to an outside AI provider.

04

Tested on real searches

We build an evaluation set from searches your producers actually ran, and the assistant is scored on it before every change.

// example projects

Priced examples for this service.

Each one is hypothetical, labelled as such, and priced live from our rate card. Open any of them in the estimator and make it yours.

Example projectWeb appAIOpen-source model

Archive search assistant

Hypothetical. Not a client, not a result.

// The problem

A broadcaster holds decades of programmes and stories, and producers spend hours hunting for a clip because the archive only knows titles and dates.

// What we would build

Every programme transcribed and indexed, with an assistant that finds the moment a subject was discussed and links to the timecode. It runs on open-source models on GPUs the broadcaster controls, because unreleased material and rights-restricted footage should not go to an outside provider.

// What is in it

  • Transcripts for the back catalogue and every new programme
  • Search by subject, person and place, down to the timecode
  • Answers that link to the clip and show the rights notes beside it
  • Speaker and language detection for English and French
  • An evaluation set from real producer searches, run before every change

// Stack

  • Whisper (open-source)
  • Llama or Qwen (open-source)
  • vLLM
  • PostgreSQL with pgvector
  • DigitalOcean GPU

// estimate

Build
≈ US$109,000 to US$166,000, delivered within 21 weeksCAD 154,900 to 236,700
Hosting
≈ US$4,530 a monthCAD 6,450 a month
Support
≈ US$1,000 a monthCAD 1,425 a month

Prices in your currency are estimates from today's Bank of Canada rate. All invoicing is in CAD or USD.

AI route
Open-source model
Model running cost
≈ US$4,190 a monthCAD 5,960 a month

Timeline by milestone

Discovery
2.9 to 3.4 weeks
Specification and evaluation plan
0.4 weeks
Design approved
2.4 to 3.9 weeks
Core features
1.4 to 1.9 weeks
Full build
0.9 to 1.4 weeks
Working pilot
3.4 to 5.4 weeks
Testing and fixes
0.9 to 1.9 weeks
Launch
0.9 to 1.4 weeks
Production
0.9 to 1.4 weeks

How it is paid

Deposit 20%
CAD 30,920 to 47,280
Discovery 2%
CAD 3,092 to 4,728
Specification and evaluation plan 10.3%
CAD 15,923.80 to 24,349.20
Design approved 2%
CAD 3,092 to 4,728
Core features 6%
CAD 9,276 to 14,184
Full build 4%
CAD 6,184 to 9,456
Working pilot 31.1%
CAD 48,080.60 to 73,520.40
Testing and fixes 2%
CAD 3,092 to 4,728
Launch 2%
CAD 3,092 to 4,728
Production 10.6%
CAD 16,387.60 to 25,058.40
Holdback, 30 days after launch (10%)
CAD 15,460 to 23,640
GPU time for training and testing, at cost
CAD 300. Billed up front, at cost, outside the milestones

// questions

Questions we get about this.

How long does it take to transcribe a whole archive?

It depends on the size of the back catalogue and how many GPUs run at once. The estimate prices the GPU time, and new programmes are transcribed as they air.

Could archive search use a frontier model instead?

Yes, for material already published, and it is cheaper to start. For unreleased or restricted footage, an open-source model on GPUs you control keeps it private; we price both.

Does it handle French and English programmes?

Yes. Transcription and search work in both, and results show which language the clip is in.

// next

Not quite your project?

Tell us what you have in mind. We will come back to you with a range and the questions that would narrow it. Or book a call and talk it through.

AI archive search for media on open-source models | Atheron Network Labs