What is the use of Spark sql's batch physics plan BatchScanExec
This article will explain in detail the use of BatchScanExec, a batch physics plan for Spark sql. The editor thinks it is very practical, so I share it with you for reference. I hope you can get something after reading this article.
BatchScanExec is the physical plan of the batch class, and the corresponding logical plan is DataSourceV2Relation, which is Datasource.
Its input parameter is the Scan class, which has two important methods, one to get the partition list information, and the other to get the reader factory.
Override lazy val partitions: Seq [InputPartition] = batch.planInputPartitions () override lazy val readerFactory: PartitionReaderFactory = batch.createReaderFactory () override lazy val inputRDD: RDD [InternalRow] = {new DataSourceRDD (sparkContext, partitions, readerFactory, supportsColumnar)}
The planInputPartitions method gets the partition list; the createReaderFactory gets the partition reader factory, both of which determine a DataSourceRDD to act as an inputRDD object.
For traditional DataSource classes, you can use them simply by implementing the scan subclass of the corresponding data source.
The corresponding physical plans for StreamingDataSourceV2Relation are MicroBatchScanExec and ContinuousScanExec, so Scan is not needed at this time. Instead, two flow definition classes, MicroBatchStream and ContinuousStream, are used.
This is the end of this article on "what is the use of Spark sql's batch physics plan BatchScanExec". I hope the above content can be of some help to you, so that you can learn more knowledge. if you think the article is good, please share it out for more people to see.