[英]Ignoring (but not removing) NA in a dplyr group_by function
在上一篇文章中,我得到了幫助,根據多個其他變量的條件更改單個變量。
然而,由於我在分組變量中有多個缺失值,因此出現了進一步的復雜化。 下面是一個示例數據框:
df2 <- data.frame(
ID = c(101:110),
Name = c("AA", "BB", "AA", "DD", "EE", "FF", "AA", "GG", "DD", "HH"),
Age = c(1, 56, 1, 72, 12, 43, 1, 32, 72, 99),
Gender = c("F", "M", "F", NA , NA, "M", "F", "M", NA, "M"),
Group = c(1, 2, 1, 2, 1, 4, 1, 3, 2, 4),
Date = seq(from = as.Date("2019-01-01"), to = as.Date("2019-01-10"), by = 'day'),
Order = c("re-do", "first", "first", "first", "re-do", "first", "re-do", "first", "re-do", "first"),
Site = c(2, 54, 2, 522, 3, 490, 2, 23, 522, 21)
)
>df2
ID Name Age Gender Group Date Order Site
1 101 AA 1 F 1 2019-01-01 re-do 2
2 102 BB 56 M 2 2019-01-02 first 54
3 103 AA 1 F 1 2019-01-03 first 2
4 104 DD 72 <NA> 2 2019-01-04 first 522
5 105 EE 12 <NA> 1 2019-01-05 re-do 3
6 106 FF 43 M 4 2019-01-06 first 490
7 107 AA 1 F 1 2019-01-07 re-do 2
8 108 GG 32 M 3 2019-01-08 first 23
9 109 DD 72 <NA> 2 2019-01-09 re-do 522
10 110 HH 99 M 4 2019-01-10 first 21
我有一個功能,根據名稱,年齡,性別和組進行分組,然后根據日期和順序列更改ID:
library(dplyr)
df2 %>%
group_by(Name, Age, Gender, Group, Site) %>%
mutate(first_date = ifelse(Order == "first",
Date,
Date[Order == "first"])) %>%
mutate(ID = ifelse(n() > 1 & Date >= first_date,
ID[Order == "first"],
ID)) %>%
select(-first_date)
但是,我遇到的問題是,NA值仍然匹配並使用(請參閱下面第4行和第9行中復制的ID值):
ID Name Age Gender Group Date Order Site
<int> <fct> <dbl> <fct> <dbl> <date> <fct> <dbl>
1 101 AA 1 F 1 2019-01-01 re-do 2
2 102 BB 56 M 2 2019-01-02 first 54
3 103 AA 1 F 1 2019-01-03 first 2
4 104 DD 72 NA 2 2019-01-04 first 522
5 105 EE 12 NA 1 2019-01-05 re-do 3
6 106 FF 43 M 4 2019-01-06 first 490
7 103 AA 1 F 1 2019-01-07 re-do 2
8 108 GG 32 M 3 2019-01-08 first 23
9 104 DD 72 NA 2 2019-01-09 re-do 522
10 110 HH 99 M 4 2019-01-10 first 21
Warning messages:
1: Factor `Gender` contains implicit NA, consider using `forcats::fct_explicit_na`
2: Factor `Gender` contains implicit NA, consider using `forcats::fct_explicit_na`
3: Factor `Gender` contains implicit NA, consider using `forcats::fct_explicit_na`
4: Factor `Gender` contains implicit NA, consider using `forcats::fct_explicit_na`
我想要發生的是帶有NA的行被忽略但沒有被刪除(這是我在管道中使用na_omit()
設法獲得的唯一結果),所以它看起來像這樣:
ID Name Age Gender Group Date Order Site
1 101 AA 1 F 1 2019-01-01 re-do 2
2 102 BB 56 M 2 2019-01-02 first 54
3 103 AA 1 F 1 2019-01-03 first 2
4 104 DD 72 <NA> 2 2019-01-04 first 522
5 105 EE 12 <NA> 1 2019-01-05 re-do 3
6 106 FF 43 M 4 2019-01-06 first 490
7 103 AA 1 F 1 2019-01-07 re-do 2
8 108 GG 32 M 3 2019-01-08 first 23
9 109 DD 72 <NA> 2 2019-01-09 re-do 522
10 110 HH 99 M 4 2019-01-10 first 21
我認為對Gender
列中的NA
值進行額外檢查應該可以解決問題嗎?
library(dplyr)
df2 %>%
group_by(Name, Age, Gender, Group, Site) %>%
mutate(first_date = ifelse(Order == "first",
Date,
Date[Order == "first"]),
ID = ifelse(n() > 1 & Date >= first_date & !is.na(Gender),
ID[Order == "first"],
ID)) %>%
select(-first_date)
# ID Name Age Gender Group Date Order Site
# <int> <fct> <dbl> <fct> <dbl> <date> <fct> <dbl>
# 1 101 AA 1 F 1 2019-01-01 re-do 2
# 2 102 BB 56 M 2 2019-01-02 first 54
# 3 103 AA 1 F 1 2019-01-03 first 2
# 4 104 DD 72 NA 2 2019-01-04 first 522
# 5 105 EE 12 NA 1 2019-01-05 re-do 3
# 6 106 FF 43 M 4 2019-01-06 first 490
# 7 103 AA 1 F 1 2019-01-07 re-do 2
# 8 108 GG 32 M 3 2019-01-08 first 23
# 9 109 DD 72 NA 2 2019-01-09 re-do 522
#10 110 HH 99 M 4 2019-01-10 first 21
聲明:本站的技術帖子網頁,遵循CC BY-SA 4.0協議,如果您需要轉載,請注明本站網址或者原文地址。任何問題請咨詢:yoyou2525@163.com.