Developer Builds Custom Binary Format That Shrinks JSON by 80%
Packing Binary Is Fun
After a Twitter debate about JSON not being for 'Real™ developers,' one coder decided to create a binary schema language called jBin. What started as a simple file write turned into a full-fledged format with varints, field numbers, and a custom schema syntax. The result cuts JSON payloads by 80% and proves that packing binary can actually be fun.
I ended up building an entire binary schema language that can shrink JSON payloads by 80%.
- trashb
I'm missing some important parts that I would argue any (binary)format needs.
- a magic header
- a version number
- a crc or data corruption check
Additionally I would argue one would be better off writing a custom text parser instead of parsing a binary format in this case, this is similar to the debate about unixlike config vs windows regedit. I would prefer something other then json but sill readable as txt.
Even the json example provided at the bottom can be minified from 418 characters to 166 by replacing the field names with single characters and removing the spaces. Almost all of the savings in this format come from not including the the field names and having a position dependent layout. You can choose to have delimiters or arrange for a byte to indicate the type(+size) for example the following string encodes the example data almost (85bytes vs 80bytes) as efficient but is still readable and supports utf-8 interpretation.
123456789;LeroyJenkins;60;alliance;p,100,200,300;i,999,1,1;i,45,100,0;a,s,120;a,a,45;
You may optimize it further by allowing recurring entries and allowing assumed values from a defined default and only sending delta's can be dependent on the type of data you are expecting.
{ "itemId": 999, "quantity": 1, "isSoulbound": false }
i,999,1,0
could become:
default = { "itemId": 0, "quantity": 1, "isSoulbound": false }
{ "itemId": 999}
i,999
- IvanK_net
There already exists a "binary analogy" of JSON called Protocol Buffers or Kiwi. They are pretty simple (a parser can be made in 2 kB of code). Wouold be great if the author compared his reinvented wheel with existing wheels :D
- Skwid
One of my favourite yak shaving adventures in a previous job was writing a parse in place UBJSON decoder for ~1MB of data on a device with about as much free memory. Fast (enough) access by key, binary search with a few shortcuts for the mostly numeric payload.
It built an index of maybe 30 bytes to speed things along, but also to let it keep working on the tail of the old message whilst the new one was overwriting it's start.
Should the system design have required all this of a device with 4MB of memory? Probably not. But it worked, and I had a great time