Git commits carry kilobytes of overhead before you add a single file
How big is a Git commit?
Git stores objects as zlib-compressed data, so commit size depends on file compressibility and eventual packfile deduplication. Experiments with loose objects show a fresh repo costs about 64 KB, a 250-file commit adds 47 KB, and a 3-byte change in a 50-file directory adds 17 KB. Larger commits compress efficiently: a 16.4 MB source file stores in 1.6 MB, and a 580 KB binary in 300 KB.
So that’s 1.6 Mb to commit a 16.4 Mb source file! I expected the text to compress nicely, but that’s really impressive.
- cocoto
> The du options are: s to summarize and b to display bytes.
Small nitpick: Use long options and your code samples become self explanatory!
- nathanpankon
I played with making a git alternative that was streamlined according to workflow. I haven't completely abandoned it, but my core idea was that you can use ast path to compress the data further.
Like you said git is really efficient, and even though I went in with some criticism because of the work(mess) that is chromium and the deep hacking I did there to fix stuff between versions without forcing a complete rebuild
I came to the conclusion that as a compression alternative, my approach wasn't worth it. Git did a better job!
Anyway here's the attempt: https://github.com/pankon/gat
- jamesblonde
Git was designed with assumptions about local file storage which make it challenge when storing data in a distributed file system. We built our distributed filesystem, HopsFS, to store data in S3. To support high performance git in FUSE, our writes hit remote NVMe disk(s) and asynchronously sync to S3 (you can replicate on NVMes or just do failure recovery). Previously, we used remote NVMe disks as a write-through cache, but that doesn't work with git, S3 latency is too high.
- lucasoshiro
In newer versions of Git you check this by using the new command `git repo structure`[1]
[1] https://git-scm.com/docs/git-repo#Documentation/git-repo.txt...
- CodesInChaos
Does git support better compression algorithms than deflate, like zstd or brotli?