So, firstly I have some inputs like this:
A:<phone1,phone2>,<location1>,<email1>
B:<phone1>,<location2>,<email1,email2>
I'd like to use Pyspark.rdd.map()function to loop every time in the row and turn them into key-value pairs like this:
phone1: A:<phone1,phone2>,<location1>,<email1>
phone1: B:<phone1>,<location2>,<email1,email2>
phone2: A:<phone1,phone2>,<location1>,<email1>
location1: A:<phone1,phone2>,<location1>,<email1>
location2: B:<phone1>,<location2>,<email1,email2>
email1: A:<phone1,phone2>,<location1>,<email1>
email1: B:<phone1>,<location2>,<email1,email2>
email2: B:<phone1>,<location2>,<email1,email2>
In my previous attempts, I tried to add a loop onto the lambda function inside of the map function, but it didn't support it. Is there any other way?
scala> val rdd = sc.parallelize(Seq("A:<phone1,phone2>,<location1>,<email1>", "B:<phone1>,<location2>,<email1,email2>"))
scala> rdd.foreach(println)
A:<phone1,phone2>,<location1>,<email1>
B:<phone1>,<location2>,<email1,email2>
scala> case class dataclass(c0:String, c1:String)
scala> val df = rdd.map(x => x.split(":")).map(y => dataclass(y(0), y(1))).toDF
scala> df.show(false)
+---+------------------------------------+
|c0 |c1 |
+---+------------------------------------+
|A |<phone1,phone2>,<location1>,<email1>|
|B |<phone1>,<location2>,<email1,email2>|
+---+------------------------------------+
scala> val df1 = df.withColumn("tempCol",regexp_replace(regexp_replace(col("c1"), "<", ""),">", ""))
.withColumn("tempCol", explode(split(col("tempCol"), ",")))
.withColumn("out", concat(col("tempCol"), lit(":"), col("c0"), lit(":"), col("c1")))
.drop("c0", "c1", "tempCol")
scala> df1.show(false)
+------------------------------------------------+
|out |
+------------------------------------------------+
|phone1:A:<phone1,phone2>,<location1>,<email1> |
|phone2:A:<phone1,phone2>,<location1>,<email1> |
|location1:A:<phone1,phone2>,<location1>,<email1>|
|email1:A:<phone1,phone2>,<location1>,<email1> |
|phone1:B:<phone1>,<location2>,<email1,email2> |
|location2:B:<phone1>,<location2>,<email1,email2>|
|email1:B:<phone1>,<location2>,<email1,email2> |
|email2:B:<phone1>,<location2>,<email1,email2> |
+------------------------------------------------+
scala> val rdd2 = df1.rdd.map(_(0))
scala> rdd2.foreach(println)
phone1:A:<phone1,phone2>,<location1>,<email1>
phone2:A:<phone1,phone2>,<location1>,<email1>
location1:A:<phone1,phone2>,<location1>,<email1>
email1:A:<phone1,phone2>,<location1>,<email1>
phone1:B:<phone1>,<location2>,<email1,email2>
location2:B:<phone1>,<location2>,<email1,email2>
email1:B:<phone1>,<location2>,<email1,email2>
email2:B:<phone1>,<location2>,<email1,email2>
The technical post webpages of this site follow the CC BY-SA 4.0 protocol. If you need to reprint, please indicate the site URL or the original address.Any question please contact:yoyou2525@163.com.