您好,登錄后才能下訂單哦!
spark性能優化要注意哪幾點,很多新手對此不是很清楚,為了幫助大家解決這個難題,下面小編將為大家詳細講解,有這方面需求的人可以來學習下,希望你能有所收獲。
默認用的是java序列化,但是會很慢,第二種很快,但是不一定能實現所有序列化 第二種,有些自定義類你需要在代碼中注冊(Kryo)
def main(args: Array[String]) { val sparkConf = new SparkConf() val sc = new SparkContext(sparkConf) val names = Array[String]("G304","G305","G306") val genders = Array[String]("male","female") val addresses = Array[String]("beijing","shenzhen","wenzhou","hangzhou") val infos = new ArrayBuffer[Info]() for (i<-1 to 1000000){ val name = names(Random.nextInt(3)) val gender = genders(Random.nextInt(2)) val address = addresses((Random.nextInt(4))) infos += Info(name, gender, address) } val rdd = sc.parallelize(infos) rdd.persist(StorageLevel.MEMORY_ONLY_SER) rdd.count() // rdd.persist(StorageLevel.MEMORY_ONLY) sc.stop() } case class Info(name:String, gender:String, address:String) }
def main(args: Array[String]) { val sparkConf = new SparkConf() sparkConf.registerKryoClasses(Array(classOf[Info])) val sc = new SparkContext(sparkConf) val names = Array[String]("G304","G305","G306") val genders = Array[String]("male","female") val addresses = Array[String]("beijing","shenzhen","wenzhou","hangzhou") val infos = new ArrayBuffer[Info]() for (i<-1 to 1000000){ val name = names(Random.nextInt(3)) val gender = genders(Random.nextInt(2)) val address = addresses((Random.nextInt(4))) infos += Info(name, gender, address) } val rdd = sc.parallelize(infos) rdd.persist(StorageLevel.MEMORY_ONLY_SER) rdd.count() // rdd.persist(StorageLevel.MEMORY_ONLY_SER) sc.stop()
sparkConf.registerKryoClasses(Array(classOf[Info]))
看完上述內容是否對您有幫助呢?如果還想對相關知識有進一步的了解或閱讀更多相關文章,請關注億速云行業資訊頻道,感謝您對億速云的支持。
免責聲明:本站發布的內容(圖片、視頻和文字)以原創、轉載和分享為主,文章觀點不代表本網站立場,如果涉及侵權請聯系站長郵箱:is@yisu.com進行舉報,并提供相關證據,一經查實,將立刻刪除涉嫌侵權內容。