簡體   English   中英

用於java.util.UUID的Spark數據集的不同行為

[英]Different behavior of Spark Dataset for java.util.UUID

我正在使用Spark 2.0.0並使用SparkSession創建Dataset 當我在createDataFrame方法中使用java.util.UUID ,它可以正常工作。 但是,當我將java.util.UUID作為Javabean java.util.UUID的字段並且使用此Javabean創建數據集時,它給了我scala.MatchError 請參閱下面的代碼和控制台日志。 誰能告訴我這是怎么回事,以及如何在Javabean類中使用UUID創建Dataset 謝謝。

UUIDTest.java

public class UUIDTest { 
  public static void main(String[] args) {
     SparkSession spark = SparkSession
              .builder()
              .appName("UUIDTest")
              .config("spark.sql.warehouse.dir", "/file:C:/temp")
              .master("local[2]")
              .getOrCreate();

     System.out.println("====> Create Dataset using UUID"); 

     //Working
     List<UUID> uuids = Arrays.asList(UUID.randomUUID(),UUID.randomUUID());      
     Dataset<Row> uuidSet = spark.createDataFrame(uuids, UUID.class);        
     uuidSet.show();

     System.out.println("====> Create Dataset using UserUUID"); 

     //Not Working
     List<UserUUID> userUuids = Arrays.asList(new UserUUID(UUID.randomUUID()),new UserUUID(UUID.randomUUID()));
     Dataset<Row> userUuidSet = spark.createDataFrame(userUuids, UserUUID.class);//Exception at this line        
     userUuidSet.show();    

     spark.stop();
   }
}

UserUUID.java

public class UserUUID implements Serializable{

private UUID uuid;

public UserUUID() {
}

public UserUUID(UUID uuid) {
    this.uuid = uuid;
}

public UUID getUuid() {
    return uuid;
}

public void setUuid(UUID uuid) {
    this.uuid = uuid;
  }
}

控制台輸出

16/08/26 22:49:23 INFO SharedState: Warehouse path is '/file:C:/temp'.
====> Create Dataset using UUID
16/08/26 22:49:26 INFO CodeGenerator: Code generated in 248.230818 ms
16/08/26 22:49:26 INFO CodeGenerator: Code generated in 10.550477 ms
+--------------------+-------------------+
|leastSignificantBits|mostSignificantBits|
+--------------------+-------------------+
|-6786538026241948655|5045373365275148508|
|-9161219066266259673|6040751881536491488|
+--------------------+-------------------+

====> Create Dataset using UserUUID
Exception in thread "main" scala.MatchError: 4fa3941c-f312-4031-a61b-01f2acef751b (of class java.util.UUID)
at org.apache.spark.sql.catalyst.CatalystTypeConverters$StructConverter.toCatalystImpl(CatalystTypeConverters.scala:256)
at org.apache.spark.sql.catalyst.CatalystTypeConverters$StructConverter.toCatalystImpl(CatalystTypeConverters.scala:251)
at org.apache.spark.sql.catalyst.CatalystTypeConverters$CatalystTypeConverter.toCatalyst(CatalystTypeConverters.scala:103)
at org.apache.spark.sql.catalyst.CatalystTypeConverters$$anonfun$createToCatalystConverter$2.apply(CatalystTypeConverters.scala:403)
at org.apache.spark.sql.SQLContext$$anonfun$beansToRows$1$$anonfun$apply$1.apply(SQLContext.scala:1106)
at org.apache.spark.sql.SQLContext$$anonfun$beansToRows$1$$anonfun$apply$1.apply(SQLContext.scala:1106)
at scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)
at scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)
at scala.collection.IndexedSeqOptimized$class.foreach(IndexedSeqOptimized.scala:33)
at scala.collection.mutable.ArrayOps$ofRef.foreach(ArrayOps.scala:186)
at scala.collection.TraversableLike$class.map(TraversableLike.scala:234)
at scala.collection.mutable.ArrayOps$ofRef.map(ArrayOps.scala:186)
at org.apache.spark.sql.SQLContext$$anonfun$beansToRows$1.apply(SQLContext.scala:1106)
at org.apache.spark.sql.SQLContext$$anonfun$beansToRows$1.apply(SQLContext.scala:1104)
at scala.collection.Iterator$$anon$11.next(Iterator.scala:409)
at scala.collection.Iterator$class.toStream(Iterator.scala:1322)
at scala.collection.AbstractIterator.toStream(Iterator.scala:1336)
at scala.collection.TraversableOnce$class.toSeq(TraversableOnce.scala:298)
at scala.collection.AbstractIterator.toSeq(Iterator.scala:1336)
at org.apache.spark.sql.SparkSession.createDataFrame(SparkSession.scala:373)
at com.UUIDTest.main(UUIDTest.java:30)
16/08/26 22:49:26 INFO SparkContext: Invoking stop() from shutdown hook

在經歷了許多嘗試使其在遇到此問題時能夠正常工作的努力之后,我發現的唯一解決方案是使用list<text>而不是list<uuid>並在要使用該方法使用UUID時在Java級別進行映射: UUID.fromString(uuidStr)

暫無
暫無

聲明:本站的技術帖子網頁,遵循CC BY-SA 4.0協議,如果您需要轉載,請注明本站網址或者原文地址。任何問題請咨詢:yoyou2525@163.com.

 
粵ICP備18138465號  © 2020-2024 STACKOOM.COM