OpenLake - High-Performance Storage for AI Workloads

Show HN: We Beat MLPerf: Modern Storage for KV Offload and LLM Training

OpenLake - High-Performance Storage for AI Workloads

OpenLake is a high-performance storage system that leverages io_uring and GPUDirect Storage to deliver low-latency, high-throughput I/O for AI training and inference. In the MLPerf Storage v3.0 benchmark, OpenLake achieved the highest read and write bandwidth among comparable S3 submissions for Llama 3.1 8B checkpointing, with write bandwidth 1.98x faster than the next competitor. Its Infinity Core I/O Engine uses asynchronous I/O and fine-grained coalescing to minimize GPU idle time during checkpoints and accelerate recovery, making it ideal for large-scale LLM training. OpenLake is open-source and available on GitHub.

We are thrilled with OpenLake’s participation in MLPerf Storage v3.0, including the new checkpointing workload, to help the community advance AI training.