Class VortexIo

java.lang.Object
dev.vortex.spark.io.VortexIo
All Implemented Interfaces:
Serializable

public final class VortexIo extends Object implements Serializable
The Hadoop configuration this Spark job reaches storage with, and the read settings that go with it.

File contents are read and written through Hadoop streams, so the connector sees the same schemes and credential providers as Spark's file index and commit protocol. Reads open a HadoopReadable; writes go through HadoopWritable on the task path Spark's commit protocol assigns.

Built on the driver and shipped to executors, so the configuration travels with it.

See Also:
  • Field Summary

    Fields
    Modifier and Type
    Field
    Description
    static final String
    Option bounding how many concurrent read upcalls the native reader issues against one file.
  • Method Summary

    Modifier and Type
    Method
    Description
    static VortexIo
    create(VortexOptions options, org.apache.hadoop.conf.Configuration hadoopConf)
    Captures hadoopConf and the read settings found in the format options.
    static VortexIo
    Hadoop I/O over a default configuration, for callers that have no Spark session to draw one from.
    org.apache.hadoop.conf.Configuration
     
    dev.vortex.io.NativeReadable
    Opens a byte source for file, stating it only if the listing that produced it did not report a size.
    int
    Bound on concurrent read upcalls per file; zero keeps the native default.

    Methods inherited from class java.lang.Object

    clone, equals, finalize, getClass, hashCode, notify, notifyAll, toString, wait, wait, wait
  • Field Details

    • READ_CONCURRENCY_OPTION

      public static final String READ_CONCURRENCY_OPTION
      Option bounding how many concurrent read upcalls the native reader issues against one file. Each in-flight upcall leases one pooled Hadoop input stream. Zero keeps the native default.
      See Also:
  • Method Details

    • create

      public static VortexIo create(VortexOptions options, org.apache.hadoop.conf.Configuration hadoopConf)
      Captures hadoopConf and the read settings found in the format options.
    • defaults

      public static VortexIo defaults()
      Hadoop I/O over a default configuration, for callers that have no Spark session to draw one from.
    • hadoopConf

      public org.apache.hadoop.conf.Configuration hadoopConf()
    • readConcurrency

      public int readConcurrency()
      Bound on concurrent read upcalls per file; zero keeps the native default.
    • openReadable

      public dev.vortex.io.NativeReadable openReadable(VortexFile file)
      Opens a byte source for file, stating it only if the listing that produced it did not report a size. The caller owns the result and must close it once the scan built on it is done.

      The readers skip listed files without VortexFile.EXTENSION before they get here. The check stops any other path from opening such a file as Vortex data.

      Throws:
      IllegalArgumentException - if the path does not carry VortexFile.EXTENSION