Class VortexFooterReader
-
Field Summary
FieldsModifier and TypeFieldDescriptionstatic final StringOption bounding how many footers are read at once.static final StringOption bounding how many footers scan statistics may read;0removes the bound.static final StringOption turning schema merging off, leaving the first file's schema to stand for the dataset. -
Method Summary
Modifier and TypeMethodDescriptionstatic OptionalLongestimatedRowCount(VortexFile file, VortexIo io, VortexOptions options) Returns the footer row count for one file, exact or estimated.static OptionalLongexactRowCount(VortexFile file, VortexIo io, VortexOptions options) Returns the footer row count for one file, and only when the footer states it exactly.static org.apache.spark.sql.types.StructTypeinferSchema(List<org.apache.hadoop.fs.FileStatus> files, VortexIo io, VortexOptions options) Infers the data schema of a Vortex dataset, or returns null when the listing holds no files at all.static OptionalLongsumRowCounts(List<VortexFile> files, VortexIo io, VortexOptions options) Sums footer row counts in a bounded pool.
-
Field Details
-
MAX_FILES_OPTION
Option bounding how many footers scan statistics may read;0removes the bound.- See Also:
-
MERGE_SCHEMA_OPTION
Option turning schema merging off, leaving the first file's schema to stand for the dataset.- See Also:
-
FOOTER_PARALLELISM_OPTION
Option bounding how many footers are read at once.- See Also:
-
-
Method Details
-
inferSchema
public static org.apache.spark.sql.types.StructType inferSchema(List<org.apache.hadoop.fs.FileStatus> files, VortexIo io, VortexOptions options) Infers the data schema of a Vortex dataset, or returns null when the listing holds no files at all.Listed files without
VortexFile.EXTENSIONare not part of the dataset and are skipped.Every file's footer is read and the schemas are merged, so a column only some files carry is still part of the dataset. A field missing from a file is nullable in the result, and the reader fills it with nulls for that file's rows. Set "vortex.mergeSchema" to false to read one footer and let the first file's schema stand for the dataset; a dataset of uniform files then costs one footer read instead of one per file.
- Throws:
IllegalArgumentException- if the listing holds files but none of them is a Vortex file, or if two files give a field types that cannot be merged
-
estimatedRowCount
Returns the footer row count for one file, exact or estimated.Only for Spark scan statistics, which are an estimate by contract. Anything that must answer a query with this number needs
exactRowCount(dev.vortex.spark.io.VortexFile, dev.vortex.spark.io.VortexIo, dev.vortex.spark.VortexOptions). -
exactRowCount
Returns the footer row count for one file, and only when the footer states it exactly.COUNT(*)pushdown answers the query from this number instead of reading the file, so an estimate would be returned to the user as fact. -
sumRowCounts
Sums footer row counts in a bounded pool.Returns empty if any footer has no count at all, or if the dataset holds more files than "vortex.stats.maxFiles" allows. Each footer costs a read against storage on the driver, so a large dataset would otherwise pay for the whole listing before the job starts.
-