將包Elastic（嵌套列表？）的R輸出轉換為data.frame或JSON

Question

我正在使用R和包'elastic'來查詢包含JSON格式的twitter數據的彈性搜索數據庫。 查詢工作正常，我得到了我期望的輸出內容（out）。

class(out) 
[1] "list"

$ hits $ hits返回

> out$hits$hits
[[1]]
[[1]]$`_index`
[1] "twitter_all_geo-2014-11-01"

[[1]]$`_type`
[1] "ctweet"

[[1]]$`_id`
[1] "ubicity-twitter-160f0964-6fc7-43ef-af2a-0e1b8c8184c7"

[[1]]$`_version`
[1] 1

[[1]]$`_score`
[1] 2.10757

[[1]]$`_source`
[[1]]$`_source`$id
[1] "528330489049120770"

[[1]]$`_source`$created_at
[1] "2014-10-31T23:39:39+0000"

[[1]]$`_source`$user
[[1]]$`_source`$user$name
[1] "afterlifetemis"


[[1]]$`_source`$place
[[1]]$`_source`$place$geo_point 
[[1]]$`_source`$place$geo_point[[1]]
[1] 30.4529

[[1]]$`_source`$place$geo_point[[2]]
[1] 50.61104


[[1]]$`_source`$place$city
[1] "Ukraine"

[[1]]$`_source`$place$country
[1] "Ukraine"

[[1]]$`_source`$place$country_code
[1] "UA"

[[1]]$`_source`$msg
[[1]]$`_source`$msg$text
[1] "u had one job artemis\none"

[[1]]$`_source`$msg$lang
[1] "EN"

[[1]]$`_source`$msg$hash_tags
list()

[[2]]
[[2]]$`_index`
[1] "twitter_all_geo-2014-11-01"

[[2]]$`_type`
[1] "ctweet"
...
...

基本上我想將數據保存為.csv文件，所以我輸入了

> write.csv(out$hits$hits,'out.csv')
Error in data.frame(text = "u had one job artemis\none", lang = "EN",   : arguments imply differing number of rows: 1, 0

我認為有必要將其轉換為data.frame，所以我試過：

> df <- ldply (out, data.frame)

data.frame中的錯誤（text =“你有一個工作artemis \\ none”，lang =“EN”，：參數意味着不同的行數：1,0

（我嘗試了其他幾個，樂觀主義者，嘗試太像這一個:)

> t(sapply(out$hits$hits, '[', 1:max(sapply(out$hits$hits, length))))
  _index                       _type    _id                                                        _version _score  _source
[1,] "twitter_all_geo-2014-11-01" "ctweet" "ubicity-twitter-160f0964-6fc7-43ef-af2a-0e1b8c8184c7" 1        2.10757 List,5 
[2,] "twitter_all_geo-2014-11-01" "ctweet" "ubicity-twitter-ba071fff-cafb-4d3f-947d-13c934905c1b" 1        2.10757 List,5 
[3,] "twitter_all_geo-2014-11-01" "ctweet" "ubicity-twitter-dd64af32-4d59-4008-a3db-74471ad269d1" 1        2.10757 List,5 
[4,] "twitter_all_geo-2014-11-01" "ctweet" "ubicity-twitter-4ba0d3d0-642d-4f9f-aaf9-c55929c35dc4" 1        2.10757 List,5 
[5,] "twitter_all_geo-2014-11-01" "ctweet" "ubicity-twitter-d7b8cbbc-87b3-44b5-8c9c-91c7b62f1458" 1        2.10757 List,5 
[6,] "twitter_all_geo-2014-11-01" "ctweet" "ubicity-twitter-76353a7c-44c9-4863-a59d-adb16716ca18" 1        2.10757 List,5 
[7,] "twitter_all_geo-2014-11-01" "ctweet" "ubicity-twitter-2aec0798-9918-4b66-9b2a-ef5a4d1f3711" 1        2.10757 List,5 
[8,] "twitter_all_geo-2014-11-01" "ctweet" "ubicity-twitter-c9e7637d-358a-40ee-a06c-85af04c22191" 1        2.10757 List,5 
[9,] "twitter_all_geo-2014-11-01" "ctweet" "ubicity-twitter-8928c1ef-f46a-4682-99c4-4dbc55270b03" 1        2.10757 List,5 
[10,] "twitter_all_geo-2014-11-01" "ctweet" "ubicity-twitter-d6b19975-b310-46c4-af11-af56971b7c4b" 1        2.10757 List,5

並且在開始時它看起來很好，但實際的推文消息不再在矩陣中

我很樂觀，並認為可能首先將其轉換為JSON（使用RJSON）

toJSON（out）toJSON（out）出錯：無法轉義字符串。 字符串不是utf8

最后我有一個列表，無法保存，無法轉換為JSON，data.frame或data.table（因為它不統一）。 有沒有人可以給我一個提示a）將其轉換為JSON或如何將列表保存為.csv文件或將其放入data.frame？

非常感謝，我想我不明白。

-Tobias

Answer 1

我認為unlist()和matrix()可以完成這項工作。

一個例子將所述Search() -返回out到數據幀：

# get the first 3 hits from elasticsearch store
out <- Search(index="shakespeare", size=3)

# (optional) verify that all hits expand to the same length
# (should be true for data intended to be in a table format)
stopifnot(
    sapply(
        out$hits$hits, 
        function(x) {!(length(unlist(x)) - length(unlist(out$hits$hits[[1]])))}
    )
)

# count number of columns, use unlist() to convert 
# nested lists to a vector, use the first hit as proxy
nColumns <- length(unlist(out$hits$hits[[1]]))

# fetch column names ... as above
nNames <- names(unlist(out$hits$hits[[1]]))

# unlist all hits and convert to matrix with ncol Columns, don't forget byrow=TRUE!
df <- data.frame(matrix(unlist(out$hits$hits), ncol=nColumns, byrow=TRUE))

# setting the column names
names(df) <- nNames

# do whatever you want with df
print(df)

干杯!

Answer 2

您可以在R中使用“jqr”包。例如： -

datacsv<-jq(out,".hits.hits[] | @csv")

它會將您的數據保存為csv格式，在“jqr”的幫助下，您還可以grep您想要的字段。

將包Elastic（嵌套列表？）的R輸出轉換為data.frame或JSON

問題描述

2 個解決方案

解決方案1
2 2015-04-23 08:36:38

解決方案2
1 2016-09-17 14:00:14

將包Elastic（嵌套列表？）的R輸出轉換為data.frame或JSON

問題描述

2 個解決方案

解決方案1 2 2015-04-23 08:36:38

解決方案2 1 2016-09-17 14:00:14

解決方案1
2 2015-04-23 08:36:38

解決方案2
1 2016-09-17 14:00:14