Spark2.3 RDD之 distinct 源码浅谈

最新推荐文章于 2023-03-21 15:21:23 发布

原创最新推荐文章于 2023-03-21 15:21:23 发布 · 2.2k 阅读

1 ·

CC 4.0 BY-SA版权

文章标签：

#distinct #spark #scala

spark 专栏收录该内容

19 篇文章

订阅专栏

本文介绍了Spark中RDD的distinct操作实现原理。通过源码分析，详细解释了如何利用map和reduceByKey算子来完成元素去重，并展示了如何通过设置期望的分区数来优化去重过程。

distinct 源码：

/**

 * Return a new RDD containing the distinct elements in this RDD.
 */
def distinct(numPartitions: Int)(implicit ord: Ordering[T] = null): RDD[T] = withScope {
  map(x => (x, null)).reduceByKey((x, y) => x, numPartitions).map(_._1)
}

/**
 * Return a new RDD containing the distinct elements in this RDD.
 */
def distinct(): RDD[T] = withScope {
  distinct(partitions.length)
}

这个去重算子比较直观了，其实就是把map 和 reduceByKey 封装了一下。distinct有一个可选参数numPartitions，这个参数是你期望的分区数。

例子：

object DistinctTest extends App {

  val sparkConf = new SparkConf().
    setAppName("TreeAggregateTest")
    .setMaster("local[6]")

  val spark = SparkSession
    .builder()
    .config(sparkConf)
    .getOrCreate()

  val value: RDD[Int] = spark.sparkContext.parallelize(List(1, 2, 3, 5, 8, 9), 3)
  println(value.distinct(1).getNumPartitions)
}

最后的结果分区被重置为1。