使用R＆dplyr计算多列中的出现次数

Question

This should be a simple solution...I just can't wrap my head around this. 这应该是一个简单的解决方案...我只是无法解决这个问题。 I'd like to count the occurrences of a factor across multiple columns of a data frame. 我想计算一个数据帧的多个列中一个因素的出现。 There're 13 columns range from abx.1 > abx.13 and a huge number of rows. 从abx.1> abx.13一共有13列，并且有大量行。

Sample data frame: 样本数据框：

library(dplyr)
 abx.1 <- c('Amoxil', 'Cipro', 'Moxiflox', 'Pip-tazo')
 start.1 <- c('2012-01-01', '2012-02-01', '2013-01-01', '2014-01-01')
 abx.2 <- c('Pip-tazo', 'Ampicillin', 'Amoxil', NA)
 start.2 <- c('2012-01-01', '2012-02-01', '2013-01-01', NA)
 abx.3 <- c('Ampicillin', 'Amoxil', NA, NA)
 start.3 <- c('2012-01-01', '2012-02-01', NA,NA)
 worksheet <-data.frame (abx.1, start.1, abx.2, start.2, abx.3, start.3)

Result I'd like: 结果我想要：

name count 名字计数
Amoxil 3 阿莫西尔3
Ampicillin 2 氨苄青霉素2
Pip-tazo 2 ip唑2
Cipro 1 Cipro 1
Moxiflox 1 Moxiflox 1

I've tried : 我试过了：

worksheet %>% group_by (abx.1, abx.2, abx.3) %>% summarise(count = n())

This doesn't give me my desired output. 这没有给我我想要的输出。 Any thoughts would be greatly appreciated. 任何想法将不胜感激。

Answer 1

If you want a dplyr solution, I'd suggest combining it with tidyr in order to convert your data to a long format first 如果您需要dplyr解决方案，建议您将其与tidyr结合使用，以便首先将数据转换为长格式

library(tidyr)
worksheet %>%
  select(starts_with("abx")) %>%
  gather(key, value, na.rm = TRUE) %>%
  count(value)

# Source: local data frame [5 x 2]
# 
#        value n
# 1     Amoxil 3
# 2 Ampicillin 2
# 3      Cipro 1
# 4   Moxiflox 1
# 5   Pip-tazo 2

Alternatively, with base R, it's just 或者，使用底数R

as.data.frame(table(unlist(worksheet[grep("^abx", names(worksheet))])))
#         Var1 Freq
# 1     Amoxil    3
# 2      Cipro    1
# 3   Moxiflox    1
# 4   Pip-tazo    2
# 5 Ampicillin    2

使用R＆dplyr计算多列中的出现次数

问题描述

1 个解决方案

解决方案1
3 已采纳 2015-07-14 21:01:53

使用R＆dplyr计算多列中的出现次数

问题描述

1 个解决方案

解决方案1 3 已采纳 2015-07-14 21:01:53

解决方案1
3 已采纳 2015-07-14 21:01:53