Measuring Format Effects on Scan Efficiency: A Minimal, Reproducible Methodology

Authors

  • Ankit Joshi

Keywords:

Parquet, ORC, CSV, NDJSON, predicate pushdown, scan efficiency, columnar storage, DuckDB, reproducibility

Abstract

File and record formats are often treated as interchangeable substrates for analytics, yet query cost is strongly shaped by format-level mechanisms such as predicate pushdown, column projection, statistics, and parse overhead. This article presents a compact, reproducible methodology for measuring how format choice influences scan efficiency and tail latency on a single node. Using one dataset in multiple encodings (Parquet [3], ORC [8], CSV, NDJSON [5]), the method varies selectivity, column width, and stripe/row-group size and compression, reporting P50/P95 latency, bytes scanned, and row-group skip rate while detecting pushdown hits and misses. The contribution is a format-centric measurement recipe that others can re-run, extend, and compare across environments, with public artifacts [9] enabling full reproduction. In the refreshed runs (TPC-H SF1, 15 repetitions), sorted Parquet with 8 MB row-groups and ZSTD scanned ~23 MB at ≈1% selectivity versus 0.77 GB (CSV) and 2.35 GB (NDJSON), with P95 latency of 2–3 ms versus 226–239 ms; skip rate ≈94%. ORC (sorted, 8 MB stripes, ZSTD) scanned approximately 0.28 GB - fewer bytes than CSV — yet produced a P95 latency of approximately 850 ms, exceeding all other formats, including CSV. This result isolates a critical ecosystem constraint: columnar byte savings are only realized as latency savings when the analytical engine has a native scanner with predicate pushdown support.

Downloads

Download data is not yet available.

References

M. Stonebraker, D. J. Abadi, A. Batkin, X. Chen, M. Cherniack, and M. Ferreira, “C-Store: A column-oriented DBMS,” in Proc. VLDB, 2005, pp. 553–564.

S. Melnik, A. Gubarev, J. Long, G. Romer, S. Shivakumar, M. Tolton, and T. Vassilakis, “Dremel: Interactive analysis of web-scale datasets,” in Proc. VLDB, 2010, pp. 330–339.

Apache Software Foundation, “Apache Parquet file format,” 2024. [Online]. Available: https://parquet.apache.org/docs/file-format/

Apache Software Foundation, “Apache Avro schema resolution,” 2024. [Online]. Available: https://avro.apache.org/docs/current/specification/

JSON Lines, “NDJSON format specification,” 2024. [Online]. Available: https://jsonlines.org/

J. Jain, A. Deshpande, and A. Gupta, “Analyzing and comparing lakehouse storage systems,” in Proc. CIDR, 2023.

H. Pirk, M. Rigger, and M. Kersten, “DuckDB: An analytical database system in-process,” in Proc. VLDB (systems overview), 2017.

Apache Software Foundation, “Apache ORC format specification,” 2024. [Online]. Available: https://orc.apache.org/docs/

A. Joshi, “FormatBench (companion repository),” 2024. [Online]. Available: https://github.com/ankitjoshi14/FormatBench

Downloads

Published

31.07.2026

How to Cite

Ankit Joshi. (2026). Measuring Format Effects on Scan Efficiency: A Minimal, Reproducible Methodology. International Journal of Intelligent Systems and Applications in Engineering, 14(1s), 2104 –. Retrieved from https://ijisae.org/index.php/IJISAE/article/view/8485

Issue

Section

Research Article