Index
All Classes and Interfaces|All Packages|Constant Field Values|Serialized Form
A
- ArrowUtils - Class in dev.vortex.spark
-
Utility class for converting Arrow types to Spark SQL data types.
- asCaseSensitiveMap() - Method in class dev.vortex.spark.VortexOptions
-
The options with the keys as given, for the case-sensitive world of Hadoop configuration.
B
- BATCH_SIZE_OPTION - Static variable in class dev.vortex.spark.write.VortexOutputWriter
-
Option sizing the row batch converted to Arrow and handed to Vortex at a time.
- buildColumnarReader(PartitionedFile) - Method in class dev.vortex.spark.read.VortexAggregateReaderFactory
- buildColumnarReader(PartitionedFile) - Method in class dev.vortex.spark.read.VortexPartitionReaderFactory
- buildReader(PartitionedFile) - Method in class dev.vortex.spark.read.VortexAggregateReaderFactory
- buildReader(PartitionedFile) - Method in class dev.vortex.spark.read.VortexPartitionReaderFactory
C
- close() - Method in class dev.vortex.spark.io.HadoopWritable
- close() - Method in class dev.vortex.spark.read.VortexArrowColumnVector
-
No-op: the underlying Arrow
ValueVectors are owned by theArrowReaderthat produced this batch and are released when that reader is closed. - close() - Method in class dev.vortex.spark.read.VortexPartitionReader
- close() - Method in class dev.vortex.spark.write.VortexOutputWriter
- convert(Filter, StructType) - Static method in class dev.vortex.spark.read.SparkFilterToVortexExpression
-
Converts a filter, returning empty when its operator, column, or literal is unsupported.
- convert(StructType) - Static method in class dev.vortex.spark.write.SparkToArrowSchema
-
Converts a Spark StructType schema to an Arrow Schema.
- create(VortexOptions, Configuration) - Static method in class dev.vortex.spark.io.VortexIo
-
Captures
hadoopConfand the read settings found in the format options. - create(Configuration, String) - Static method in class dev.vortex.spark.io.HadoopWritable
-
Creates (or overwrites) the file at
path, along with any missing parent directories.
D
- defaults() - Static method in class dev.vortex.spark.io.VortexIo
-
Hadoop I/O over a default configuration, for callers that have no Spark session to draw one from.
- dev.vortex.spark - package dev.vortex.spark
- dev.vortex.spark.io - package dev.vortex.spark.io
- dev.vortex.spark.read - package dev.vortex.spark.read
- dev.vortex.spark.write - package dev.vortex.spark.write
E
- empty() - Static method in class dev.vortex.spark.VortexOptions
-
No options at all, for callers with no Spark read or write to draw them from.
- equals(Object) - Method in class dev.vortex.spark.VortexOptions
- estimatedRowCount(VortexFile, VortexIo, VortexOptions) - Static method in class dev.vortex.spark.read.VortexFooterReader
-
Returns the footer row count for one file, exact or estimated.
- exactRowCount(VortexFile, VortexIo, VortexOptions) - Static method in class dev.vortex.spark.read.VortexFooterReader
-
Returns the footer row count for one file, and only when the footer states it exactly.
- EXTENSION - Static variable in class dev.vortex.spark.io.VortexFile
-
The extension every Vortex data file must carry.
F
- flush() - Method in class dev.vortex.spark.io.HadoopWritable
- FOOTER_PARALLELISM_OPTION - Static variable in class dev.vortex.spark.read.VortexFooterReader
-
Option bounding how many footers are read at once.
- fromArrowField(Field) - Static method in class dev.vortex.spark.ArrowUtils
-
Converts an Arrow Field to a Spark SQL DataType.
- fromArrowType(ArrowType) - Static method in class dev.vortex.spark.ArrowUtils
-
Converts an Arrow type to a Spark SQL DataType.
G
- get() - Method in class dev.vortex.spark.read.VortexPartitionReader
- get() - Method in interface dev.vortex.spark.VortexSessionProvider
-
Construct (or return a cached)
Session. - get() - Static method in class dev.vortex.spark.VortexSparkSession
-
Returns the default JVM-wide session, creating it on first use.
- get(VortexOptions) - Static method in class dev.vortex.spark.VortexSparkSession
-
Resolve the session to use for a given set of Spark format options.
- get(String) - Method in class dev.vortex.spark.VortexOptions
-
The value set for
keyunder any casing. - get(String, String) - Method in class dev.vortex.spark.VortexOptions
-
The value set for
keyunder any casing, orfallbackwhen it is not set. - getArray(int) - Method in class dev.vortex.spark.read.VortexArrowColumnVector
-
Returns the array value at the specified row.
- getBinary(int) - Method in class dev.vortex.spark.read.VortexArrowColumnVector
-
Returns the binary data (byte array) at the specified row.
- getBoolean(int) - Method in class dev.vortex.spark.read.VortexArrowColumnVector
-
Returns the boolean value at the specified row.
- getBoolean(String, boolean) - Method in class dev.vortex.spark.VortexOptions
-
The value set for
keyas a boolean, orfallbackwhen it is not set. - getByte(int) - Method in class dev.vortex.spark.read.VortexArrowColumnVector
-
Returns the byte value at the specified row.
- getChild(int) - Method in class dev.vortex.spark.read.VortexArrowColumnVector
-
Returns the child column at the specified ordinal.
- getDecimal(int, int, int) - Method in class dev.vortex.spark.read.VortexArrowColumnVector
-
Returns the decimal value at the specified row with the given precision and scale.
- getDouble(int) - Method in class dev.vortex.spark.read.VortexArrowColumnVector
-
Returns the double value at the specified row.
- getFileExtension(TaskAttemptContext) - Method in class dev.vortex.spark.write.VortexOutputWriterFactory
- getFloat(int) - Method in class dev.vortex.spark.read.VortexArrowColumnVector
-
Returns the float value at the specified row.
- getInt(int) - Method in class dev.vortex.spark.read.VortexArrowColumnVector
-
Returns the int value at the specified row.
- getInt(String, int) - Method in class dev.vortex.spark.VortexOptions
-
The value set for
keyas an integer, orfallbackwhen it is not set. - getLong(int) - Method in class dev.vortex.spark.read.VortexArrowColumnVector
-
Returns the long value at the specified row.
- getMap(int) - Method in class dev.vortex.spark.read.VortexArrowColumnVector
-
Returns the map value at the specified row.
- getShort(int) - Method in class dev.vortex.spark.read.VortexArrowColumnVector
-
Returns the short value at the specified row.
- getUTF8String(int) - Method in class dev.vortex.spark.read.VortexArrowColumnVector
-
Returns the UTF8String value at the specified row.
- getValueVector() - Method in class dev.vortex.spark.read.VortexArrowColumnVector
-
Returns the underlying Apache Arrow ValueVector wrapped by this column vector.
H
- hadoopConf() - Method in class dev.vortex.spark.io.VortexIo
- HadoopReadable - Class in dev.vortex.spark.io
-
A
NativeReadablethat serves Vortex's reads from a HadoopFileSystem. - HadoopWritable - Class in dev.vortex.spark.io
-
A
NativeWritablethat streams a Vortex file into a HadoopFileSystem. - hashCode() - Method in class dev.vortex.spark.VortexOptions
- hasNull() - Method in class dev.vortex.spark.read.VortexArrowColumnVector
-
Returns whether this column contains any null values.
- hasVortexExtension(String) - Static method in class dev.vortex.spark.io.VortexFile
-
Returns whether
pathnames a Vortex data file.
I
- inferSchema(List<FileStatus>, VortexIo, VortexOptions) - Static method in class dev.vortex.spark.read.VortexFooterReader
-
Infers the data schema of a Vortex dataset, or returns null when the listing holds no files at all.
- isNullAt(int) - Method in class dev.vortex.spark.read.VortexArrowColumnVector
-
Returns whether the value at the specified row is null.
- isPushable(Filter, StructType) - Static method in class dev.vortex.spark.read.SparkFilterToVortexExpression
-
Returns whether the complete filter can be evaluated by Vortex for
dataSchema.
L
- length() - Method in class dev.vortex.spark.io.VortexFile
M
- MAX_FILES_OPTION - Static variable in class dev.vortex.spark.read.VortexFooterReader
-
Option bounding how many footers scan statistics may read;
0removes the bound. - MERGE_SCHEMA_OPTION - Static variable in class dev.vortex.spark.read.VortexFooterReader
-
Option turning schema merging off, leaving the first file's schema to stand for the dataset.
N
- newInstance(String, StructType, TaskAttemptContext) - Method in class dev.vortex.spark.write.VortexOutputWriterFactory
- next() - Method in class dev.vortex.spark.read.VortexPartitionReader
- numNulls() - Method in class dev.vortex.spark.read.VortexArrowColumnVector
-
Returns the total number of null values in this column.
O
- of(Map<String, String>) - Static method in class dev.vortex.spark.VortexOptions
-
Captures
optionsas given. - open(Configuration, String) - Static method in class dev.vortex.spark.io.HadoopReadable
-
Opens a readable over
path, stating the file for its length. - open(Configuration, String, long) - Static method in class dev.vortex.spark.io.HadoopReadable
-
Opens a readable over
path. - openReadable(VortexFile) - Method in class dev.vortex.spark.io.VortexIo
-
Opens a byte source for
file, stating it only if the listing that produced it did not report a size. - openStream() - Method in class dev.vortex.spark.io.HadoopReadable
- options() - Method in class dev.vortex.spark.read.VortexAggregateReaderFactory
- options() - Method in class dev.vortex.spark.read.VortexPartitionReaderFactory
P
- path() - Method in class dev.vortex.spark.io.VortexFile
- path() - Method in class dev.vortex.spark.write.VortexOutputWriter
- PROVIDER_OPTION - Static variable in class dev.vortex.spark.VortexSparkSession
-
Options key used to select a
VortexSessionProviderby class name.
R
- READ_CONCURRENCY_OPTION - Static variable in class dev.vortex.spark.io.VortexIo
-
Option bounding how many concurrent read upcalls the native reader issues against one file.
- readConcurrency() - Method in class dev.vortex.spark.io.VortexIo
-
Bound on concurrent read upcalls per file; zero keeps the native default.
- requireVortexExtension(String) - Static method in class dev.vortex.spark.io.VortexFile
-
Fails unless
pathnames a Vortex data file.
S
- setDefault(Session) - Static method in class dev.vortex.spark.VortexSparkSession
-
Replace the default session.
- SparkFilterToVortexExpression - Class in dev.vortex.spark.read
-
Converts Spark's stable V1 file filters into Vortex expressions.
- SparkToArrowSchema - Class in dev.vortex.spark.write
-
Utility class for converting Spark SQL schemas to Arrow schemas.
- sumRowCounts(List<VortexFile>, VortexIo, VortexOptions) - Static method in class dev.vortex.spark.read.VortexFooterReader
-
Sums footer row counts in a bounded pool.
- supportColumnarReads(InputPartition) - Method in class dev.vortex.spark.read.VortexAggregateReaderFactory
- supportColumnarReads(InputPartition) - Method in class dev.vortex.spark.read.VortexPartitionReaderFactory
- supportsDataType(DataType) - Static method in class dev.vortex.spark.write.SparkToArrowSchema
-
Returns whether Vortex's Spark writer can represent this Spark SQL type.
T
- toString() - Method in class dev.vortex.spark.VortexOptions
U
- UNKNOWN_LENGTH - Static variable in class dev.vortex.spark.io.VortexFile
-
Size of a file nothing has stat'ed yet.
- unsized(String) - Static method in class dev.vortex.spark.io.VortexFile
-
A file whose size is not known yet.
V
- VortexAggregateReaderFactory - Class in dev.vortex.spark.read
-
Produces one footer-backed partial COUNT(*) row per Vortex file.
- VortexAggregateReaderFactory(FileSourceOptions, VortexIo, VortexOptions, StructType, StructType, Aggregation) - Constructor for class dev.vortex.spark.read.VortexAggregateReaderFactory
- VortexArrowColumnVector - Class in dev.vortex.spark.read
-
Spark ColumnVector implementation that wraps Apache Arrow vectors from Vortex data.
- VortexArrowColumnVector(ValueVector) - Constructor for class dev.vortex.spark.read.VortexArrowColumnVector
-
Creates a new VortexArrowColumnVector wrapping the specified Arrow ValueVector.
- VortexFile - Class in dev.vortex.spark.io
-
A Vortex file and, when a listing already reported it, its size on storage.
- VortexFile(String, long) - Constructor for class dev.vortex.spark.io.VortexFile
- VortexFooterReader - Class in dev.vortex.spark.read
-
Reads schema and row-count metadata from Vortex file footers.
- VortexIo - Class in dev.vortex.spark.io
-
The Hadoop configuration this Spark job reaches storage with, and the read settings that go with it.
- VortexOptions - Class in dev.vortex.spark
-
The format options of one Spark read or write, resolved without regard to key case.
- VortexOutputWriter - Class in dev.vortex.spark.write
-
Writes Spark InternalRow data to a Vortex file.
- VortexOutputWriter(String, StructType, VortexOptions, NativeWritable) - Constructor for class dev.vortex.spark.write.VortexOutputWriter
-
Creates a writer for the task path assigned by Spark's commit protocol.
- VortexOutputWriterFactory - Class in dev.vortex.spark.write
-
Creates Vortex output writers at paths assigned by Spark's file commit protocol.
- VortexOutputWriterFactory(StructType, VortexOptions) - Constructor for class dev.vortex.spark.write.VortexOutputWriterFactory
- VortexPartitionReader - Class in dev.vortex.spark.read
-
Columnar reader over one Spark
PartitionedFile. - VortexPartitionReader(PartitionedFile, StructType, StructType, StructType, VortexIo, VortexOptions, Filter[], boolean) - Constructor for class dev.vortex.spark.read.VortexPartitionReader
- VortexPartitionReaderFactory - Class in dev.vortex.spark.read
-
Produces one Vortex reader for each file selected by Spark's file index.
- VortexPartitionReaderFactory(FileSourceOptions, VortexIo, VortexOptions, StructType, StructType, StructType, Filter[], boolean) - Constructor for class dev.vortex.spark.read.VortexPartitionReaderFactory
- VortexSessionProvider - Interface in dev.vortex.spark
-
User hook for supplying a custom
Sessionto Vortex Spark readers and writers. - VortexSparkSession - Class in dev.vortex.spark
-
JVM-wide holder for one or more Vortex
Sessions used by Spark readers and writers.
W
- WORKER_THREADS_OPTION - Static variable in class dev.vortex.spark.VortexSparkSession
-
Options key sizing the JVM-wide pool of background threads that drive Vortex futures.
- write(byte[], int, int) - Method in class dev.vortex.spark.io.HadoopWritable
- write(InternalRow) - Method in class dev.vortex.spark.write.VortexOutputWriter
-
Writes a single row to the Vortex file.
- WRITE_BATCH_SIZE_OPTION - Static variable in class dev.vortex.spark.write.VortexOutputWriter
-
Option sizing the write batch for Vortex alone, overriding "batch.size".
All Classes and Interfaces|All Packages|Constant Field Values|Serialized Form