I’ve spent most of my career at the intersection of web development and security — and the standard architecture for adding AI to an application makes the security half of my brain itch. The default pattern is:
- take your data
- serialize it
- POST it to someone else’s computer
Every document you embed, every query your users type, every transcript you process — shipped off-premises to a third party. You’re now subject to their retention policies, their breach surface, and any potential changes to their terms-of-service.
For plenty of workloads, that trade is justifiable. For others, it’s somewhere between uncomfortable and disqualifying. This post is the case for the alternative: keep your data and your models local.
The data-governance argument
The moment user content leaves your infrastructure, your compliance story gets a new chapter. It’s also one you don’t have the luxury of writing yourself. Questions you now answer with someone else’s documentation:
- Is the vendor retaining request payloads? For how long?
- Does sending this data cross a regulatory boundary — HIPAA, GDPR data-residency, attorney-client material, PCI-scoped records?
- Is your data eligible for the vendor’s training pipelines? Can they unilaterally change their ToS provisions to change your answer?
- When the vendor has an incident, is your customers’ content in the blast radius?
Local inference removes the concern. Text embedded in-process never traverses a network. There is no payload for a vendor to retain because there is no request to an outside party. The audit conversation shifts from “review the subprocessor’s SOC 2” to “the data never left the box.”
Enterprise customers love that statement.
The economics argument
Cloud AI pricing is a meter that runs on every call. Per-token charges feel negligible in development but compound brutally in production. Think through the different things that burn tokens:
- re-embedding a corpus after a model upgrade
- embedding every search query at peak traffic
- transcribing an archive of hours-long audio
Local models invert the cost curve — you pay once in hardware and electricity, and the marginal cost of inference is zero. For high-volume, steady workloads (search queries are the canonical example), the crossover arrives fast.
The latency and reliability argument
An embedding call to a cloud API costs you a network round trip — tens to hundreds of milliseconds of pure overhead. Plus a dependency on someone else’s uptime, rate limits, and deprecation calendar.
In-process inference has no network in the loop at all. The model is memory-mapped into your process; latency is whatever the math takes.
Your search feature works when the vendor is down, because you are your own vendor.
“But don’t local models lag the frontier?”
For open-ended chat over all human knowledge — yes, frontier cloud models are better. But look at what production AI features actually do: embed text, search vectors, transcribe audio, extract structured data, summarize documents. These are narrow, well-bounded tasks, and small open-weight models have gotten remarkably good at them. Qwen3-Embedding-0.6B holds its own near the top of embedding leaderboards. Whisper transcription has been commodity-grade for years. A quantized sub-1B model runs a classification or extraction task all day on a CPU.
The honest framing is local by default and cloud where capability genuinely demands it.
Route the bounded tasks to hardware you control; spend cloud tokens (and cloud risk) on the tasks that legitimately need a frontier model. Most applications discover the bounded tasks are 90% of the volume.
Why PHP, and why now
The unfortunate gap is that the local-AI ecosystem grew up speaking Python. PHP applications wanting local inference often need to stand up a Python sidecar or run a separate inference server. With these requirements, most teams shrug and reach for the cloud API after all.
The GGUF model format and the llama.cpp / whisper.cpp ecosystem quietly removed the hard part: world-class inference engines now exist as embeddable native libraries with no Python anywhere in their dependency chain. What was missing was the bridge into the PHP runtime.
That bridge is what my next four posts are about: ext-infer for LLM inference and embeddings, ext-turbovec for in-process vector search, ext-whisper for audio transcription, and the contracts package that ties them into the broader PHP ecosystem.
In-process, open source, MIT-licensed — and your data never leaves the machine.
