Stop Reparsing the Same Big JSON Documents: Persist Them as Indexed Binary Instead
A service loads a 40 MB product catalog from disk every time a worker starts. It parses the JSON into objects, builds a map from SKU to product, and answers price lookups until the next deploy. The parse costs seconds of startup. The object tree often costs several times the file’s size in memory, and every worker holds its own copy. A typical request touches two fields of one product. Multiply that by every deploy, every autoscale event and every cold start. Nothing here is broken. The format was built for exchange and it’s being used as a database.
The waste is real but narrow. It shows up with big, read-mostly documents that get parsed again and again. A catalog loaded by every worker is the classic case; a config blob read on every request and an API dump queried over and over are the same problem. It doesn’t pay for small payloads or write-heavy data, or when a fast parser plus an in-memory cache is already enough. On a phone the same cost lands in launch time, one of the metrics that matter on mobile, and in a static site build that pulls a large remote file it lands in build time (see fetching a remote JSON API at build time).
The product worth building is specific: a schemaless, persistent, memory-mappable file with a path index. You build it once from JSON. After that, get("/items/48213/price") follows offsets through a few pages of the file and decodes one number. It parses nothing, because the parsing happened at build time. Open the file from ten processes and the operating system keeps one copy of its pages in the page cache, so memory stops growing with worker count.
The Pieces Already Exist
Postgres has had jsonb since 9.4: parsed binary JSON with indexes, inside a database server. That’s the right answer if the data already lives in Postgres and the wrong one if the document is a file next to your process. SQLite added its own binary format, JSONB, in version 3.45.0 (January 2024). It removes the text parsing and it’s the right tool when documents are rows. Its containers are sequences with sizes in the headers, so skipping a value is cheap, but finding a key in a very wide object still means walking its siblings.
simdjson parses at gigabytes per second, and its On-Demand API skips what you don’t touch. At those speeds a 40 MB file costs tens of milliseconds to scan, which is why a parser plus a cache beats any new format for load-once workloads. BSON, Amazon Ion, CBOR and MessagePack are binary JSON-like encodings, but they’re encodings: built to be compact and quick to serialize, not to jump to a path. FlatBuffers and Cap’n Proto give you zero-copy access if you write a schema first. FlatBuffers also has a schemaless sibling, FlexBuffers, which is the nearest existing thing to this idea, so any first version should be measured against it too.
Before building any of this, try changing the shape. A 40 MB catalog is usually an array of records, and an array of records is a table. Keeping API responses in SQLite with an index on json_extract(body, '$.sku') answers “price for SKU 48213” without loading anything. The format in this post is for what’s left: documents where the tree matters, where you can’t or won’t explode them into rows, or where lookups are deep and irregular. If the data arrives as a stream of small objects instead of one document, a SQL engine for JSON streams fits better. If the document is simply bigger than anyone needs, trimming responses at the proxy is cheaper than any of this.
A Tape, a String Table and a Path Index
The layout follows simdjson’s tape in spirit: a flat array of typed 8-byte words, containers that know where they end, strings in a side table. Three additions make it persistent and searchable. This is a design sketch, and the file extension and names are placeholders.
header magic, format version, flags, section offsets
tape typed words: null, bool, int64, uint64, f64, number-as-text(ref),
string(ref), array(len, position table), object(len, key index)
strings UTF-8 bytes, length-prefixed, deduplicated
keys one dictionary of object keys; objects store key ids, not text
index large objects: (key id, tape position) sorted by key bytes
large arrays: tape positions, 8 bytes per element
patches optional append-only log of JSON Patch operations
A lookup splits the JSON Pointer (RFC 6901), finds items in the root object’s sorted index with a binary search, jumps into the array’s position table at element 48213, finds price in that object and decodes one value. That’s a handful of page reads. The array table costs 8 bytes per element, so 200,000 products add 1.6 MB of offsets. Small objects skip the index and get scanned, because a scan over six keys is as fast as any index and costs no space.
let doc = Doc::open("catalog.jbin")?; // mmap and header check, no parsing
let price = doc.get("/items/48213/price")?; // a few page reads
println!("{}", price.as_f64()?); // decoded only now
for item in doc.at("/items")?.iter() { // iterate without materializing
// ...
}
Updates are where an indexed format gets expensive, because editing a value in place shifts every offset after it. So the base file stays immutable. Edits go into the append-only patch log (JSON Patch, RFC 6902, is the obvious vocabulary), readers check a small in-memory map of patched pointers before they touch the base, and a compaction pass rewrites the file when the log gets long. Read-mostly data keeps the log short, and write-heavy data was never a fit.
Where Parsers Disagree
JSON’s grammar doesn’t say how big a number can be, so parsers differ. JavaScript turns every number into a double and loses integer precision past 2^53. Python keeps arbitrary-size integers. Go’s decoder hands untyped numbers back as float64 unless you ask otherwise. A persistent format outlives the parser that wrote it, so it has to pick a policy and write it down: separate slots for int64, uint64 and double, plus a fallback that keeps the original decimal text for anything that doesn’t fit, so a 25-digit ID comes back unchanged. Keep 1.0 and 1 distinct, since the type carries information, and print doubles with the shortest form that round-trips.
The JSON spec (RFC 8259) says object names should be unique and calls the behavior with duplicates unpredictable. Real parsers keep the first, keep the last or fail. Reject duplicates at build time by default, with a flag to keep the last, because a file format shouldn’t hide a choice that changes answers. Strings are stored as UTF-8. Escapes decode at build time, a lone surrogate escape (legal in JSON text, impossible in UTF-8) is an error, and keys compare as bytes with no normalization. Normalizing would be helpful and wrong, because two keys that look the same would silently become one.
A reader that follows offsets out of a file is a parser in disguise, and a corrupt or hostile file can point anywhere. Bounds-check every offset, distrust the header’s counts, and fuzz the reader before anyone else’s files go near it. A checksum over 40 MB would make opening the file as slow as the parse it replaced, so verify on build and on request, not on every open. Version the format in the header from day one, refuse to open a newer major version, fix the byte order, and keep a corpus of old files in the test suite. A persistent format has a problem an in-memory structure never had: yesterday’s files.
Then there’s the benchmark, which decides whether any of this was worth doing. Compare against what people actually use, not a naive parser: simdjson with the parsed document kept in memory, SQLite JSONB, FlexBuffers, and a plain map built at startup. Measure cold start with an empty page cache, warm lookups, p99 per lookup, build time, file size against the original JSON and resident memory across ten processes. The format loses if the workload is one process that parses once and answers forever. If it does, say so in the README.
Version 0.1 Is Read-Only
Version 0.1 is a C or Rust library with read-only files: build from JSON, path lookup, iteration, mmap. The CLI has build, get and stats. There’s no daemon at first, no query language and no network access. A daemon adds the failure modes of a server to a library that exists to avoid doing work, and a query language is a second project that SQLite and jq have spent years on.
Build the benchmark before the format.