Class VortexOutputWriter

java.lang.Object
org.apache.spark.sql.execution.datasources.OutputWriter
dev.vortex.spark.write.VortexOutputWriter

public final class VortexOutputWriter extends org.apache.spark.sql.execution.datasources.OutputWriter
Writes Spark InternalRow data to a Vortex file.

This writer converts Spark's internal row format to Arrow vectors and writes them to a Vortex file using the Vortex writer API.

  • Field Details

    • BATCH_SIZE_OPTION

      public static final String BATCH_SIZE_OPTION
      Option sizing the row batch converted to Arrow and handed to Vortex at a time.
      See Also:
    • WRITE_BATCH_SIZE_OPTION

      public static final String WRITE_BATCH_SIZE_OPTION
      Option sizing the write batch for Vortex alone, overriding "batch.size".
      See Also:
  • Constructor Details

    • VortexOutputWriter

      public VortexOutputWriter(String filePath, org.apache.spark.sql.types.StructType schema, VortexOptions options, dev.vortex.io.NativeWritable writable)
      Creates a writer for the task path assigned by Spark's commit protocol.
      Parameters:
      filePath - the path where the Vortex file will be written
      schema - the schema of the data to write
      options - additional write options
  • Method Details

    • write

      public void write(org.apache.spark.sql.catalyst.InternalRow row)
      Writes a single row to the Vortex file.

      Rows are batched and converted to Arrow format before writing.

      Specified by:
      write in class org.apache.spark.sql.execution.datasources.OutputWriter
      Parameters:
      row - the row to write
    • close

      public void close()
      Specified by:
      close in class org.apache.spark.sql.execution.datasources.OutputWriter
    • path

      public String path()
      Specified by:
      path in class org.apache.spark.sql.execution.datasources.OutputWriter