Milford Sound

From Voice to Vectors: Building a Searchable Audio Archive in PHP

Think about the audio your organization is sitting on. Years of podcast episodes. Recorded all-hands meetings. Conference talks. Webinar archives. Sermon recordings going back a decade. Thousands of hours of institutional knowledge — none of it searchable, because audio is opaque to every tool in your stack.

With the three extensions in place, making an audio archive semantically searchable is a three-verb pipeline: transcribe → embed → search. All local, all PHP and, because audio archives deserve a privacy focus, nothing ever leaves your infrastructure.

Stage 1: Transcribe (ext-whisper)

A queue worker normalizes each file to 16kHz mono WAV (the ffmpeg one-liner from earlier) and transcribes:

<?php
$whisper = Model::load('models/ggml-small.en.gguf');
$result  = $whisper->transcribe($wavPath);

foreach ($result->segments() as $segment) {
    $store->saveSegment($episodeId, $segment);  // text + start/end times
}

Keep the segments, not just the flat transcript. Having each segment flagged individually instead of a huge blob of text will help with chunking and embedding.

An archive job like this is the canonical overnight batch: a worker grinding through a few hundred hours of audio at no marginal cost is precisely the workload where local processing embarrasses per-minute cloud transcription pricing.

Stage 2: Chunk and embed (ai-toolkit + ext-infer)

Whisper segments are short — a sentence or two — which is too fine-grained to embed individually (a five-word segment carries almost no retrievable meaning). The toolkit merges consecutive segments into ~60–90 second windows, each window inheriting the start time of its first segment:

<?php
foreach ($windows as $id => $window) {
    $vector = $embedder->embed($window->text);
    $index->addWithIds(Vectors::pack($vector->toArray()), [$id]);
}

That inherited timestamp is the magic ingredient. Every vector in the index now knows where in the audio it came from.

Stage 3: Search, and land in the right minute

<?php
$query  = $embedder->embed('what did we decide about the pricing change?');
$result = $index->search(Vectors::pack($query->toArray()), k: 5);

foreach ($result as $row) {
    $window = $store->getWindow($row['id']);
    printf("%s at %s — %s\n",
        $window->episodeTitle,
        gmdate('i:s', (int) $window->startTime),
        $window->preview(),
    );
}

The search result isn’t “this hour-long recording is relevant, good luck.” It’s “episode 42, at 14:32” — a deep link straight into the moment, with <audio>‘s media-fragment support (episode.mp3#t=872) doing the seeking for free. That’s the difference between a transcript dump and a usable archive: searchable audio is nice but navigable audio is transformative.

Because the query is embedded, instead of keyword-matched, someone searching “struggling to forgive a coworker” finds the segment where the speaker talked about grace and letting go of resentment — vocabulary the searcher never typed and the speaker never indexed.

Filtering and incremental growth

When this makes it to production, there are two details that help keep things simple:

First, the allowlist filtering from earlier: store window ids alongside episode metadata in SQL, pre-filter by show, speaker, date range, or access level, and pass the resulting id list to search() — selective filters make queries faster.

Second, online ingest: TurboQuant needs no training phase, so each newly published episode is transcribe-embed-add, searchable within minutes of upload, no index rebuilds ever.

The stack, proven

This pipeline is the last several articles distilled down into one, concrete feature: three native extensions and a toolkit composing into something none of them does alone — audio in one end, meaning-aware deep links out the other.

With zero external dependencies (or data leakage to a third party).