Popularity

6.1

Growing

Activity

9.4

Stars 654

Watchers 46

Forks 465

Last Commit about 20 hours ago

Description

ORC is a self-describing type-aware columnar file format designed for Hadoop workloads. It is optimized for large streaming reads, but with integrated support for finding required rows quickly. Storing data in a columnar format lets the reader read, decompress, and process only the values that are required for the current query. Because ORC files are type-aware, the writer chooses the most appropriate encoding for the type and builds an internal index as the file is written. Predicate pushdown uses those indexes to determine which stripes in a file need to be read for a particular query and the row indexes can narrow the search to a particular set of 10,000 rows. ORC supports the complete set of types in Hive, including the complex types: structs, lists, maps, and unions.

Code Quality Rank: L1

Programming language: Java

License: Apache License 2.0

Tags: Data Structures

Latest version: v1.6.3.rc1

Apache Orc alternatives and similar libraries

Based on the "Data structures" category.
Alternatively, view Apache Orc alternatives based on common mentions on social networks and blogs.

Protobuf

10.0 10.0 L1 Apache Orc VS Protobuf

Protocol Buffers - Google's data interchange format
Apache Thrift

9.4 8.9 L1 Apache Orc VS Apache Thrift

Apache Thrift

InfluxDB - Power Real-Time Data Analytics at Scale

Get real-time insights from all types of time series data with InfluxDB. Ingest, query, and analyze billions of data points in real-time with unbounded cardinality.

Promo www.influxdata.com

Apache Avro

8.2 9.7 L1 Apache Orc VS Apache Avro

Apache Avro is a data serialization system.
Wire

8.1 9.6 L1 Apache Orc VS Wire

gRPC and protocol buffers for Android, Kotlin, Swift and Java.
Apache Parquet

8.0 9.2 L2 Apache Orc VS Apache Parquet

Apache Parquet
SBE

7.9 8.5 L1 Apache Orc VS SBE

Simple Binary Encoding (SBE) - High Performance Message Codec
Tape

7.3 0.0 L4 Apache Orc VS Tape

A lightning fast, transactional, file-based FIFO for Android and Java.
Big Queue

5.3 0.0 Apache Orc VS Big Queue

A big, fast and persistent queue based on memory mapped file.
Persistent Collection

5.1 6.6 L4 Apache Orc VS Persistent Collection

A Persistent Java Collections Library
dexx

3.2 0.0 L3 Apache Orc VS dexx

Persistent (immutable) collections for Java and Kotlin
hjson-java

3.1 6.5 L3 Apache Orc VS hjson-java

Hjson for Java

* Code Quality Rankings and insights are calculated and provided by Lumnify.
They vary from L1 to L5 with "L5" being the highest.

Do you think we are missing an alternative of Apache Orc or a related project?

Add another 'Data structures' Library

Popular Comparisons

README

Apache ORC

ORC File Library

This project includes both a Java library and a C++ library for reading and writing the Optimized Row Columnar (ORC) file format. The C++ and Java libraries are completely independent of each other and will each read all versions of ORC files.

Releases:

Latest: Apache ORC releases
Maven Central:
Downloads: Apache ORC downloads
Release tags: Apache ORC release tags
Plan: Apache ORC future release plan

The current build status:

Main branch

Bug tracking: Apache Jira

The subdirectories are:

c++ - the c++ reader and writer
cmake_modules - the cmake modules
docker - docker scripts to build and test on various linuxes
examples - various ORC example files that are used to test compatibility
java - the java reader and writer
proto - the protocol buffer definition for the ORC metadata
site - the website and documentation
tools - the c++ tools for reading and inspecting ORC files

Building

Install java 1.8 or higher
Install maven 3.8.6 or higher
Install cmake 3.12 or higher

To build a release version with debug information:

% mkdir build
% cd build
% cmake ..
% make package
% make test-out

To build a debug version:

% mkdir build
% cd build
% cmake .. -DCMAKE_BUILD_TYPE=DEBUG
% make package
% make test-out

To build a release version without debug information:

% mkdir build
% cd build
% cmake .. -DCMAKE_BUILD_TYPE=RELEASE
% make package
% make test-out

To build only the Java library:

% cd java
% ./mvnw package

To build only the C++ library:

% mkdir build
% cd build
% cmake .. -DBUILD_JAVA=OFF
% make package
% make test-out

Apache Orc

Apache ORC - the smallest, fastest columnar storage for Hadoop workloads