重塑数据帧

Question

简单的重塑，我有以下数据：

df<-data.frame(Product=c("A","A","A","B","B","C"), Ingredients=c("Chocolate","Vanilla","Berry","Chocolate","Berry2","Vanilla"))
df
Product Ingredients
1   A   Chocolate 
2   A     Vanilla
3   A       Berry
4   B   Chocolate
5   B      Berry2
6   C     Vanilla

我想为“成分”的每个唯一值创建一个列，例如：

df2
Product Ingredient_1 Ingredient_2 Ingredient_3
A       Chocolate       Vanilla        Berry
B       Chocolate       Berry2         NULL
C       Vanilla         NULL           NULL

似乎微不足道，我尝试重塑形状，但是我一直在计数（而不是“成分”的实际值）。 想法？

Answer 1

这是使用data.table包的可能解决方案

library(data.table)
setDT(df)[, Ingredient := paste0("Ingredient_", seq_len(.N)), Product]
dcast(df, Product ~ Ingredient, value.var = "Ingredients")
#    Product Ingredient_1 Ingredient_2 Ingredient_3
# 1:       A    Chocolate      Vanilla        Berry
# 2:       B    Chocolate       Berry2           NA
# 3:       C      Vanilla           NA           NA

另外，我们可以通过性感的dplyr/tidyr组合来做到这一点

library(dplyr)
library(tidyr)
df %>% 
  group_by(Product) %>%
  mutate(Ingredient = paste0("Ingredient_", row_number())) %>%
  spread(Ingredient, Ingredients)

# Source: local data frame [3 x 4]
# 
#   Product Ingredient_1 Ingredient_2 Ingredient_3
# 1       A    Chocolate      Vanilla        Berry
# 2       B    Chocolate       Berry2           NA
# 3       C      Vanilla           NA           NA

Answer 2

本着共享替代方案的精神，这里还有两个：

选项1 ： split列，然后使用stri_list2matrix创建宽表单。

library(stringi)
x <- with(df, split(Ingredients, Product))
data.frame(Product = names(x), stri_list2matrix(x))
#   Product        X1        X2      X3
# 1       A Chocolate Chocolate Vanilla
# 2       B   Vanilla    Berry2    <NA>
# 3       C     Berry      <NA>    <NA>

选项2：使用getanID从我的“splitstackshape”包生成“.ID”一栏，然后dcast它。 “ data.table”包已加载“ splitstackshape”，因此您可以直接调用dcast.data.table进行重塑。

library(splitstackshape)
dcast.data.table(getanID(df, "Product"), 
                 Product ~ .id, value.var = "Ingredients")
#    Product         1       2     3
# 1:       A Chocolate Vanilla Berry
# 2:       B Chocolate  Berry2    NA
# 3:       C   Vanilla      NA    NA

Answer 3

底座R reshape

df$Count<-ave(rep(1,nrow(df)),df$Product,FUN=cumsum)
reshape(df,idvar="Product",timevar="Count",direction="wide",sep="_")

#  Product Ingredients_1 Ingredients_2 Ingredients_3
#1       A     Chocolate       Vanilla         Berry
#4       B     Chocolate        Berry2          <NA>
#6       C       Vanilla          <NA>          <NA>

重塑数据帧

问题描述

3 个解决方案

解决方案1
2 已采纳 2015-02-03 20:56:47

解决方案2
2 2015-02-04 04:15:58

解决方案3
1 2015-02-03 21:41:58

重塑数据帧

问题描述

3 个解决方案

解决方案1 2 已采纳 2015-02-03 20:56:47

解决方案2 2 2015-02-04 04:15:58

解决方案3 1 2015-02-03 21:41:58

解决方案1
2 已采纳 2015-02-03 20:56:47

解决方案2
2 2015-02-04 04:15:58

解决方案3
1 2015-02-03 21:41:58