Think about the audio your organization is sitting on. Years of podcast episodes. Recorded all-hands meetings. Conference talks. Webinar archives. Sermon recordings going back a decade. Thousands of hours of institutional knowledge — none of it searchable, because audio is opaque to every tool in your stack.
With the three extensions in place, making an audio archive semantically searchable is a three-verb pipeline: transcribe → embed → search. All local, all PHP and, because audio archives deserve a privacy focus, nothing ever leaves your infrastructure.
Stage 1: Transcribe (ext-whisper)
A queue worker normalizes each file to 16kHz mono WAV (the ffmpeg one-liner from earlier) and transcribes:
<?php
$whisper = Model::load('models/ggml-small.en.gguf');
$result = $whisper->transcribe($wavPath);
foreach ($result->segments() as $segment) {
$store->saveSegment($episodeId, $segment); // text + start/end times
}
Keep the segments, not just the flat transcript. Having each segment flagged individually instead of a huge blob of text will help with chunking and embedding.
An archive job like this is the canonical overnight batch: a worker grinding through a few hundred hours of audio at no marginal cost is precisely the workload where local processing embarrasses per-minute cloud transcription pricing.
Stage 2: Chunk and embed (ai-toolkit + ext-infer)
Whisper segments are short — a sentence or two — which is too fine-grained to embed individually (a five-word segment carries almost no retrievable meaning). The toolkit merges consecutive segments into ~60–90 second windows, each window inheriting the start time of its first segment:
<?php
foreach ($windows as $id => $window) {
$vector = $embedder->embed($window->text);
$index->addWithIds(Vectors::pack($vector->toArray()), [$id]);
}
That inherited timestamp is the magic ingredient. Every vector in the index now knows where in the audio it came from.
Stage 3: Search, and land in the right minute
<?php
$query = $embedder->embed('what did we decide about the pricing change?');
$result = $index->search(Vectors::pack($query->toArray()), k: 5);
foreach ($result as $row) {
$window = $store->getWindow($row['id']);
printf("%s at %s — %s\n",
$window->episodeTitle,
gmdate('i:s', (int) $window->startTime),
$window->preview(),
);
}
The search result isn’t “this hour-long recording is relevant, good luck.” It’s “episode 42, at 14:32” — a deep link straight into the moment, with <audio>‘s media-fragment support (episode.mp3#t=872) doing the seeking for free. That’s the difference between a transcript dump and a usable archive: searchable audio is nice but navigable audio is transformative.
Because the query is embedded, instead of keyword-matched, someone searching “struggling to forgive a coworker” finds the segment where the speaker talked about grace and letting go of resentment — vocabulary the searcher never typed and the speaker never indexed.
Filtering and incremental growth
When this makes it to production, there are two details that help keep things simple:
First, the allowlist filtering from earlier: store window ids alongside episode metadata in SQL, pre-filter by show, speaker, date range, or access level, and pass the resulting id list to search() — selective filters make queries faster.
Second, online ingest: TurboQuant needs no training phase, so each newly published episode is transcribe-embed-add, searchable within minutes of upload, no index rebuilds ever.
The stack, proven
This pipeline is the last several articles distilled down into one, concrete feature: three native extensions and a toolkit composing into something none of them does alone — audio in one end, meaning-aware deep links out the other.
With zero external dependencies (or data leakage to a third party).
