[英]Read a binary column in spark using java language
I have a DataFrame witch contains a Binary
column Type.我有一个 DataFrame 女巫包含一个Binary
列类型。
DataFrame: DataFrame:
+----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+
|BinaryGeometry
+----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+
|[00 00 00 00 01 03 00 00 00 01 00 00 00 11 00 00 00 04 00 F0 00 DC CC 1A C0 87 14 01 81 1E 1B 41 40 FC FF EF 00 68 AA 1A C0 BF EE 57 20 85 19 41 40 04 00 F0 00 8C 86 1A C0 CC DC 8B DC AE 1A 41 40 FF FF EF 00 44 74 1A C0 CA 9D 5D 61 10 1C 41 40 FF FF EF 00 64 63 1A C0 BF 1F 98 0B 3A 1D 41 40 FF FF EF 00 44 47 1A C0 E4 6B A0 DD CE 1D 41 40 FC FF EF 00 D8 2B 1A C0 54 E4 71 67 6D 1C 41 40 FF FF EF 00 44 1A 1A C0 BF 1F 98 0B 3A 1D 41 40 02 00 F0 00 80 0B 1A C0 0D 80 00 13 2F 23 41 40 02 00 F0 00 B0 35 1A C0 CC F6 23 F8 BD 26 41 40 04 00 F0 00 0C 43 1A C0 73 1A 44 AF 16 26 41 40 02 00 F0 00 40 5A 1A C0 FF 54 9C 7C 2D 27 41 40 02 00 F0 00 50 68 1A C0 87 6E B9 42 44 28 41 40 02 00 F0 00 00 7C 1A C0 78 2B 85 BA F5 26 41 40 FC FF EF 00 18 91 1A C0 49 96 6F 58 C6 28 41 40 02 00 F0 00 B0 BC 1A C0 91 FA 4B 0E 7F 20 41 40 04 00 F0 00 DC CC 1A C0 87 14 01 81 1E 1B 41 40] |
+----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+
I'm triying to read this column to extract the geometry format.我正在尝试阅读此专栏以提取几何格式。
from my research i found that there is a function from the geospark library which is ST_GeomFromWKB which takes a long binary as argument.根据我的研究,我发现 geospark 库中有一个 function,它是 ST_GeomFromWKB,它需要一个长二进制文件作为参数。
SO I'm doing the following code:所以我正在执行以下代码:
df.withColumn("BinaryGeometry", hex(col("BinaryGeometry")))
.withColumn("BinaryGeometry",expr("ST_GeomFromWKB(BinaryGeometry)"))
I get the following output witch is not correct:我得到以下 output 女巫不正确:
POINT(0 0)
How can I read this column to get the correct geometry value?如何阅读本专栏以获得正确的几何值?
EDIT编辑
The structute of column on MySQL MySQL上的立柱结构
And when I track the data it seems like that:当我跟踪数据时,它看起来像这样:
When I click on [GEOMETRY - 113 o]
a text file downloads.当我点击[GEOMETRY - 113 o]
时,会下载一个文本文件。
The text file contains this data:文本文件包含以下数据:
ے FPہ>Pے{ح@@€àOہZDچ4ح@@ہ¨KہإضآTح@@àمJہî=Oƒح@@û ڑLہCBsn«ح@@ے FPہ>Pے{ح@@
EDIT2编辑2
I have this table on MySQL database我在 MySQL 数据库上有这张表
+----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
|BinaryGeometry
+----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
|'POLYGON((-6.70005799736828 34.2118684058451,-6.66641236748546 34.1993751935347,-6.63139344658703 34.208461349766,-6.6135406633839 34.2192498881446,-6.5970611711964 34.2283339016717,-6.5695953508839 34.2328755410488,-6.54281617607921 34.2220887476075,-6.5256500383839 34.2283339016717,-6.51123048271984 34.2748740913813,-6.55242921318859 34.3026724029165,-6.56547547783703 34.2975672800575,-6.58813477959484 34.3060756457653,-6.60186768975109 34.314583149474,-6.62109376396984 34.3043740415805,-6.64169312920421 34.318553022848,-6.68426515068859 34.2538774367323,-6.70005799736828 34.2118684058451))',0
+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
When I load this table in Spark I get:当我在 Spark 中加载此表时,我得到:
+----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+
|BinaryGeometry
+----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+
|[00 00 00 00 01 03 00 00 00 01 00 00 00 11 00 00 00 04 00 F0 00 DC CC 1A C0 87 14 01 81 1E 1B 41 40 FC FF EF 00 68 AA 1A C0 BF EE 57 20 85 19 41 40 04 00 F0 00 8C 86 1A C0 CC DC 8B DC AE 1A 41 40 FF FF EF 00 44 74 1A C0 CA 9D 5D 61 10 1C 41 40 FF FF EF 00 64 63 1A C0 BF 1F 98 0B 3A 1D 41 40 FF FF EF 00 44 47 1A C0 E4 6B A0 DD CE 1D 41 40 FC FF EF 00 D8 2B 1A C0 54 E4 71 67 6D 1C 41 40 FF FF EF 00 44 1A 1A C0 BF 1F 98 0B 3A 1D 41 40 02 00 F0 00 80 0B 1A C0 0D 80 00 13 2F 23 41 40 02 00 F0 00 B0 35 1A C0 CC F6 23 F8 BD 26 41 40 04 00 F0 00 0C 43 1A C0 73 1A 44 AF 16 26 41 40 02 00 F0 00 40 5A 1A C0 FF 54 9C 7C 2D 27 41 40 02 00 F0 00 50 68 1A C0 87 6E B9 42 44 28 41 40 02 00 F0 00 00 7C 1A C0 78 2B 85 BA F5 26 41 40 FC FF EF 00 18 91 1A C0 49 96 6F 58 C6 28 41 40 02 00 F0 00 B0 BC 1A C0 91 FA 4B 0E 7F 20 41 40 04 00 F0 00 DC CC 1A C0 87 14 01 81 1E 1B 41 40] |
+----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+
then to get the original data 'Polygon.........' I do the following code:然后获取原始数据'Polygon............'我执行以下代码:
df.withColumn("geom",expr("ST_GeomFromWKB(BinaryGeometry)"));
But I get the following error:但我收到以下错误:
20/08/10 22:28:50 ERROR Executor: Exception in task 87.0 in stage 39.0 (TID 929) java.lang.ClassCastException: [B cannot be cast to org.apache.spark.unsafe.types.UTF8String
at org.apache.spark.sql.geosparksql.expressions.ST_GeomFromWKB.eval(Constructors.scala:174)
at org.apache.spark.sql.catalyst.expressions.GeneratedClass$SpecificUnsafeProjection.writeFields_0_39$(Unknown Source)
at org.apache.spark.sql.catalyst.expressions.GeneratedClass$SpecificUnsafeProjection.apply(Unknown Source)
at org.apache.spark.sql.catalyst.expressions.GeneratedClass$SpecificUnsafeProjection.apply(Unknown Source)
MySQL geometry types are not really stored as WKB
and as such cannot be read using the geospark's method ST_GeomFromWKB
. MySQL 几何类型并未真正存储为WKB
,因此无法使用 geospark 的方法ST_GeomFromWKB
。
https://dev.mysql.com/doc/refman/5.7/en/gis-data-formats.html https://dev.mysql.com/doc/refman/5.7/en/gis-data-formats.html
Internally, MySQL stores geometry values in a format that is not identical to either
WKT
orWKB
format.在内部,MySQL 以不同于WKT
或WKB
格式的格式存储几何值。 (Internal format is likeWKB
but with an initial 4 bytes to indicate the SRID.) (内部格式类似于WKB
,但有一个初始的 4 个字节来指示 SRID。)
Solution to that is to specify the query when loading dataframe that will parse the geometry
types into WKT
format using the builtin function ST_AsWKT
:解决方案是在加载 dataframe 时指定查询,该查询将使用内置 function ST_AsWKT
将geometry
类型解析为WKT
格式:
df = spark
.read
.format("jdbc")
.options(
Map(
"driver" -> "com.mysql.cj.jdbc.Driver",
"url" -> "jdbc:mysql://host:3306/db",
"user" -> "user",
"password" -> "password",
"dbtable" -> "(select ST_AsWKT(BinaryGeometry) as BinaryGeometry from geo_table) as t")
)
.load
.withColumn("BinaryGeometry",expr("ST_GeomFromWKT(BinaryGeometry)"))
Alternatively, the column can be extracted as WKB using或者,可以使用 WKB 将列提取为
select hex(ST_AsWKB(BinaryGeometry)) as BinaryGeometry from geo_table
and fed to ST_GeomFromWKB
method.并馈送到ST_GeomFromWKB
方法。
scala> df.printShema()
root
|-- BinaryGeometry: geometry (nullable = false)
scala> df.show()
+--------------------+
| BinaryGeometry|
+--------------------+
|POLYGON ((-7.5783...|
+--------------------+
声明:本站的技术帖子网页,遵循CC BY-SA 4.0协议,如果您需要转载,请注明本站网址或者原文地址。任何问题请咨询:yoyou2525@163.com.