Paging Through Parquet in DuckDB: file_row_number vs OFFSET
Paging Through a Parquet File in DuckDB: File_row_number or Offset?

I tested paging through large Parquet files in DuckDB to see if OFFSET causes quadratic slowdowns. Surprisingly, DuckDB rewrites OFFSET efficiently, but using file_row_number remains faster and more reliable. Crucially, I discovered that OFFSET without ORDER BY can silently drop or duplicate rows, making simple row counts useless for validating data exports.
So don't validate an export by counting rows. A count can't see this, because the drops and the duplicates cancel each other out.