我认为,目前接受的解决方案将无法工作,如果有超过1个相邻的
NA
的值
dt
.
另一种选择是,注意顺序很重要:
解决方案
dat
a dt
1 a <NA>
2 a <NA>
3 a 2017-05-01
4 a 2017-06-01
5 b <NA>
6 b 2017-08-01
7 b 2017-09-01
library(dplyr)
library(tidyr)
dat %>%
group_by(a) %>%
mutate(helper = ifelse(is.na(dt), NA, cumsum(!is.na(dt)))) %>%
fill(helper, .direction = 'up') %>%
group_by(a, helper) %>%
mutate(dt = coalesce(dt,
max(dt, na.rm = TRUE) - months(max(row_number()) - row_number()))) %>%
dplyr::select(-helper)
# A tibble: 7 x 3
# Groups: a, helper [4]
helper a dt
<int> <fct> <date>
1 1 a 2017-03-01
2 1 a 2017-04-01
3 1 a 2017-05-01
4 2 a 2017-06-01
5 1 b 2017-07-01
6 1 b 2017-08-01
7 2 b 2017-09-01
数据
dat <-structure(list(a = structure(c(1L, 1L, 1L, 1L, 2L, 2L, 2L), .Label = c("a",
"b"), class = "factor"), dt = structure(c(NA, NA, 17287, 17318,
NA, 17379, 17410), class = "Date")), .Names = c("a", "dt"), row.names = c(NA,
-7L), class = "data.frame")