Description
ORC is a self-describing type-aware columnar file format designed for Hadoop workloads. It is optimized for large streaming reads, but with integrated support for finding required rows quickly. Storing data in a columnar format lets the reader read, decompress, and process only the values that are required for the current query. Because ORC files are type-aware, the writer chooses the most appropriate encoding for the type and builds an internal index as the file is written. Predicate pushdown uses those indexes to determine which stripes in a file need to be read for a particular query and the row indexes can narrow the search to a particular set of 10,000 rows. ORC supports the complete set of types in Hive, including the complex types: structs, lists, maps, and unions.
Apache Orc alternatives and similar libraries
Based on the "Data structures" category.
Alternatively, view Apache Orc alternatives based on common mentions on social networks and blogs.
CodeRabbit: AI Code Reviews for Developers
* Code Quality Rankings and insights are calculated and provided by Lumnify.
They vary from L1 to L5 with "L5" being the highest.
Do you think we are missing an alternative of Apache Orc or a related project?
README
Apache ORC
ORC is a self-describing type-aware columnar file format designed for Hadoop workloads. It is optimized for large streaming reads, but with integrated support for finding required rows quickly. Storing data in a columnar format lets the reader read, decompress, and process only the values that are required for the current query. Because ORC files are type-aware, the writer chooses the most appropriate encoding for the type and builds an internal index as the file is written. Predicate pushdown uses those indexes to determine which stripes in a file need to be read for a particular query and the row indexes can narrow the search to a particular set of 10,000 rows. ORC supports the complete set of types in Hive, including the complex types: structs, lists, maps, and unions.
ORC File Library
This project includes both a Java library and a C++ library for reading and writing the Optimized Row Columnar (ORC) file format. The C++ and Java libraries are completely independent of each other and will each read all versions of ORC files.
Releases:
- Latest: Apache ORC releases
- Maven Central:
- Downloads: Apache ORC downloads
- Release tags: Apache ORC release tags
- Plan: Apache ORC future release plan
The current build status:
- Main branch
Bug tracking: Apache Jira
The subdirectories are:
- c++ - the c++ reader and writer
- cmake_modules - the cmake modules
- docker - docker scripts to build and test on various linuxes
- examples - various ORC example files that are used to test compatibility
- java - the java reader and writer
- proto - the protocol buffer definition for the ORC metadata
- site - the website and documentation
- tools - the c++ tools for reading and inspecting ORC files
Building
- Install java 1.8 or higher
- Install maven 3.8.6 or higher
- Install cmake 3.12 or higher
To build a release version with debug information:
% mkdir build
% cd build
% cmake ..
% make package
% make test-out
To build a debug version:
% mkdir build
% cd build
% cmake .. -DCMAKE_BUILD_TYPE=DEBUG
% make package
% make test-out
To build a release version without debug information:
% mkdir build
% cd build
% cmake .. -DCMAKE_BUILD_TYPE=RELEASE
% make package
% make test-out
To build only the Java library:
% cd java
% ./mvnw package
To build only the C++ library:
% mkdir build
% cd build
% cmake .. -DBUILD_JAVA=OFF
% make package
% make test-out