8 min read Rocky Elsalaymeh
Our vector database never ran. Retrieval was fine.
The README promised sqlite-vec embeddings. The code path threw on every call and a catch block absorbed it. Retrieval worked anyway, which is exactly why nobody looked. Here is what the audit found and what I changed.
In May I wrote that the SQLite and sqlite-vec stack for retrieval was “faster than I expected.” It was a confident line about a component that had never executed once. I praised the speed of something that was not running.
Team-X, an open-source, local-first desktop app for running AI-agent organizations, advertised sqlite-vec embeddings for retrieval. The sqlite-vec path threw an error on every call. A catch block swallowed the error and fell back to an exact cosine scan in plain code. Retrieval stayed correct, so nobody looked. At this scale, the scan was all I needed.
Here is the line, from Field Notes issue 001, in my own words: “The SQLite + sqlite-vec stack for retrieval is faster than I expected.” Whatever I was timing, it was not sqlite-vec. I am publishing this because the failure is a pattern, not a one-off, and the pattern is worth more than the embarrassment.
What did the audit actually find?
A retrieval path that could not have worked, for five reasons. Four of them were independent and individually fatal, according to the changelog entry written when the dead code was deleted on main. The fifth was a query that never used the caller’s vector.
| Defect | What the changelog records |
|---|---|
| Migration absent from the journal | 0022_sqlite_vec_integration.sql was not in migrations/meta/_journal.json, so drizzle never applied it |
| Index collision | It shared index 0022 with the journaled 0022_long_run_resume_origin.sql |
| Table defined nowhere | It INSERTed into a migration_metadata table that no schema defines |
| Extension never loaded | db/client.ts opened better-sqlite3, set three pragmas, and called no loadExtension |
| Query ignored the query vector | The query never referenced the caller’s queryVector; there was no vec0 MATCH clause |
The consequence, in the changelog’s words: every call threw no such table: embeddings_vec. The virtual table did not exist in any database Team-X had ever created.
The fourth row deserves a second look. The SQLite documentation is plain that an extension is a separate shared library: “To load it, you need to supply SQLite with the name of the file containing the shared library or DLL and an entry point to initialize the extension.” Opening a database does not do that for you. sqlite-vec itself is a small piece of software, “Written in pure C, no dependencies, runs anywhere SQLite runs,” according to its README. The tool was never the problem. I never loaded it.
The migration was also the most confident document in the repo. Its header promised “O(1) vector similarity search” and described HNSW indexing. The code that actually ran was a linear scan. The README, before the fix, listed “sqlite-vec embeddings” in its RAG bullet. The changelog’s correction is blunt: “It never had them.”
Why did nobody notice for four months?
Because the catch block worked. That is the whole story, and it is a more dangerous story than a crash.
This is the code that sat in the retrieval service before the audit:
} catch (error) {
// Fall through to brute force if similaritySearch fails
const errMsg = error instanceof Error ? error.message : String(error);
console.warn('[RAG] similaritySearch failed, falling back to brute force:', errMsg);
}
Read it as a reviewer would. It looks responsible. It catches the error, names the fallback, and logs a warning. Every instinct says this is good engineering.
It is not, for one reason. The warning fired on every retrieval. A warning that fires on every call carries no information, because it cannot tell the healthy case from the broken one. A signal has to distinguish states. This one described a single state, forever.
The user-visible output was identical either way. The changelog states that retrieval behavior is unchanged after the dead path was removed: “so RAG has always ranked by in-process cosine similarity.” The fallback converted an outage into a performance difference. Nobody alerts on a performance difference nobody measured.
The catch block that made retrieval reliable hid a feature that never worked.
This is not a new observation. In a study at USENIX OSDI 2014, Yuan and colleagues examined 198 randomly selected, user-reported failures in Cassandra, HBase, HDFS, Hadoop MapReduce, and Redis. Their finding: “We found the majority of catastrophic failures could easily have been prevented by performing simple testing on error handling code.” Error handling is where those failures lived, and mine was no exception.
The Zen of Python states the principle in two lines in PEP 20: “Errors should never pass silently. Unless explicitly silenced.” My catch block was not an explicit silence. It was a fallback that never said which path had served the request. Those are different failures, and only the second one is invisible.
The gap that hid it was mechanical. The migrations directory held 38 .sql files while the journal listed 37. The changelog calls this “the specific gap that hid this for four months.” A new test in readme-claims.test.ts now asserts that the README count, the file count, and the journal agree, and that every migration file is journaled.
Did I need a vector database at all?
No. At this scale, exact in-process cosine search served every query, and it did so correctly for the entire time the “accelerated” path was dead.
Exact search has a property that an index can never match. pgvector’s README says it directly: “By default, pgvector performs exact nearest neighbor search, which provides perfect recall.” Add an approximate index and it “trades some recall for speed,” in its words. The FAISS guidance on choosing an index agrees on the exactness point: “The only index that can guarantee exact results is the IndexFlatL2 or IndexFlatIP.”
So the real question is cost. The repo does the arithmetic in the index source: at 10k chunks and 768 dimensions, a full scan is about 7.7M multiply-adds per query, and it grows linearly with the corpus. That is the honest cost of the exact path. It is not free, and it is not frightening.
On timing, I will be precise about what I can prove. A code comment says that below 4,096 vectors “a full scan costs single-digit milliseconds and exactness is free.” That is a claim in a comment. I found no published benchmark of the scan or the index in the repo, so I am not going to give you a latency figure. If you want one for your corpus, measure it.
There is a hidden cost to the other side of the ledger. An index must be built, and it must be rebuilt when the data changes. The FAISS guide notes that for only a few searches, “the index building time will not be amortized by the search time,” and then direct computation is the most efficient option. Team-X’s per-company index is rebuilt whenever that company’s vectors change. A small, churning corpus is exactly the case where the build cost never pays back.
When does approximate search start to matter, and how does the code decide?
It matters when one company’s corpus reaches 4,096 stored vectors, and the code decides with a single comparison. The IVF index added on main is inert below that line.
Here is the decision, from the service and its branch: the floor is annOpts.minVectors ?? 4096, and the index is used only when annEnabled && rows.length >= annMinVectors. The rationale is in the option’s docs. Approximate search “would degrade answer quality for no measurable gain” on a small corpus, so the floor is the safety property.
The index is IVF, written in plain TypeScript. K-means partitions the vectors, and a query scans only the nearest partitions. I chose it over HNSW for a reason the source spells out: HNSW’s failure mode is silently degraded recall, and IVF has a property that is easy to reason about, namely “probing every cluster is brute force.” There is no native dependency either, so the index runs under plain Vitest. That was the constraint that ruled out reinstating sqlite-vec. A path I cannot test is the path that just burned me.
That equivalence matters for the theme of this post. The exact path is the ground truth, and the approximate path is a relaxation of it. If the index ever misbehaves, I can compare it against the code that has always worked.
Recall is the price. The changelog reports measured recall against brute-force ground truth:
| Partitions probed (nProbe) | Recall reported in the changelog |
|---|---|
| 1 | 62.3% |
| 6 | 93.3% |
| 10 | 97.5% |
| 32 (all) | 100% |
Two caveats, because this is a post about unverified claims. First, these are fixture numbers: the test uses 2,000 synthetic vectors over 32 clusters and probes 6. They say nothing about your embeddings. Second, I could not find those four figures in the test file. The assertion there is a floor: expect(hitCount / total).toBeGreaterThanOrEqual(0.9). The changelog says more than the test enforces, which is the same class of drift this post is about, and the same class behind docs that described a product that did not exist.
The ANN-Benchmarks paper makes the right general point. It describes a tool that “provides a standard interface for measuring the performance and quality achieved by nearest neighbor algorithms on different standard data sets.” Performance and quality. An approximate index with no recall measurement is a speed claim with the cost hidden.
How do you keep a fallback from hiding a dead feature?
Make the fallback report which path served the request, and test that the primary path is actually reachable. A fallback is only safe if the failure it absorbs is visible somewhere other than a log line.
Five concrete actions, in the order I would run them:
- List every catch that continues with a substitute result. Retrieval, caching, config loading, feature flags. For each one, ask what a human would see if the primary path failed on every call.
- Record which path answered. A counter or a field on the result, not a warning. If the fallback rate is 100%, that number should be impossible to miss.
- Write a test that forces the primary path and fails if it is not used. Then break the primary path on purpose and confirm the test fails. A green test that cannot fail proves nothing.
- Assert that your migrations and your journal agree. Team-X now checks that every migration file is journaled. It targets the specific gap that hid this one.
- Measure your corpus before buying an index. Multiply vectors by dimensions, time the exact scan on your hardware, and set an explicit floor for approximate search. Keep the exact path as the recall baseline.
None of this requires a vector database. It requires deciding what “working” means and checking it.
The corrected code and the README line that now says what the code does are in the Team-X repo. If you want the next audit as it happens, including the parts I get wrong, subscribe to Field Notes.
Frequently asked questions
Did Team-X ever use sqlite-vec for retrieval?
No. Team-X, an open-source, local-first desktop app for running AI-agent organizations, advertised sqlite-vec embeddings, but the sqlite-vec path threw on every call and a fallback caught the error. Retrieval always ranked stored chunks by exact cosine similarity computed in-process. The dead path has since been removed from the repo.
Why did the sqlite-vec failure go unnoticed in Team-X?
Team-X wrapped the sqlite-vec call in a catch block that fell through to a brute-force cosine scan. The scan returned correct results, so retrieval looked healthy. The only trace was a console warning on each retrieval. The changelog says retrieval behavior was unchanged after the dead path was deleted.
When does Team-X switch from exact to approximate vector search?
Team-X uses exact in-process cosine search by default. On main, an IVF approximate index engages only when a company's corpus reaches 4,096 stored vectors, set by the minVectors option. Below that floor retrieval stays exact, because trading recall for speed on a small corpus costs answer quality for no measurable gain.
A fallback that never reports which path served the request is a lie by omission.