How HyperCrux Was Tested: 200 Killed Writers and Not One Mismatched Record
HyperCrux makes one promise above the others. A record can be reached four ways, by key, by SQL, by its links and by its vector, and the four never disagree. A deleted record doesn’t linger in search results. A link never points at a record that isn’t there. A vector never belongs to a row that was only half written.
Promises like that are easy to make when nothing goes wrong. So the tests go looking for things going wrong, and the results are published with the code, in test/results on GitHub.

Killing the Writer
The hardest test starts a process that writes to a file as fast as it can, then kills it with SIGKILL at a random moment, 200 times over. SIGKILL gives a process no warning and no chance to tidy up. Whatever it was in the middle of simply stops.
Every transaction the writer makes touches all four handles at once. It puts a record with fields and a vector, links it to the record before it and to one of ten anchors, deletes an old record every seventh time, which takes that record’s links with it, and moves a counter. So a kill in the middle of any transaction could, in principle, leave a record without its vector, a link to a record that was half deleted, or a counter ahead of the data.
After each kill the test opens the file and compares it with what the counter says must be there. Every record that should exist, with exactly the right fields and the right vector. Every link, and not one more. No record from a transaction that didn’t commit. Then the vector handle has to find each sampled record as its own nearest neighbour, the link handle has to walk the chain correctly, and HyperCrux’s Check and SQLite’s integrity check both have to pass.
In the recorded run the writer committed 30,306 transactions across the 200 kills, ending with 25,979 records and 47,631 links. After every single kill, the file matched the last committed transaction exactly. SQLite’s transactions are what make that possible. HyperCrux’s part is to put every handle inside them: the links, the key registry and the vector checks all change in the same transaction as the row.
Sharing the File
Four processes wrote 300 transactions each to one file at the same time, linking their records to a shared hub, searching vectors and walking links between writes, while another process read the file throughout. Every record and every link arrived, and the file checked out.
Another Language, the Same Rules
HyperCrux keeps its rules as SQLite triggers inside the file, so they should hold for programs that know nothing about HyperCrux. The tests check that with Python’s own sqlite3 module and plain SQL. Python put a record with a vector and linked it to others. It was refused when it tried to put a key in the wrong table, link to a record that doesn’t exist, store a vector of the wrong size or change a key. When it deleted a record, the links went with it. Then Python and Go wrote 600 records and links into the same file at once, and the file stayed consistent.
The search was checked the same way. Python read the raw vectors, compared them with a question itself and ranked the results. It found the same closest records as HyperCrux, with distances equal to within a billionth.
Exact Means Exact
HyperCrux’s vector search compares every vector, so it should never miss a closer match. The tests hold it to that against a brute-force comparison: 2,000 vectors, 20 questions, with and without a filter, and the results had to match one for one. Thirteen more statements in plain SQL tried to break the file in different ways, from linking to nothing to storing a vector of zeros, and SQLite refused every one.
What Review Found
Before the first release, an independent review went through the code looking for ways around the promise, and it found some. A table named keys or links could lose one of its triggers to a clash of names. A REPLACE on a second unique column, such as an email address, could remove a row and leave its links behind. A schema change that one transaction rolled back could mislead the next. Each is fixed and has a test of its own, so it stays fixed.
The review also drew a line worth knowing. Inserts, updates and deletes from any program keep the file in step. Dropping or rebuilding a record table with plain SQL is a schema change, which triggers don’t see, so HyperCrux has drop and adopt commands for it, and check reports anything left over.
How Fast
The middle of three runs on a two-core cloud machine, with every committed write flushed to disk:
| Operation | Time |
|---|---|
| Put one record, committed | 0.35 ms |
| Put records with 384-value vectors, 1,000 per transaction | 51 µs a record |
| Get a record by key | 18 µs |
| Walk 1 and 3 links out, among 100,000 records with 5 links each | 43 µs and 0.51 ms |
| Nearest 10 among 10,000 vectors of 384 values | 42 ms |
| Nearest 10 among 100,000 vectors of 384 values | 0.41 s |
Single writes wait for the disk, which sets their pace more than anything else.
Search time grows with the number of vectors compared, which is the price of exact results. Up to around a hundred thousand vectors in a table, that price is small.
The race detector ran over the whole suite and found nothing. Anyone can repeat all of it with sh scripts/record-tests.sh from a clone of the repository.