代码之家  ›  专栏  ›  技术社区  ›  DeduciveR

随机或按比例为NAs分配类别值

na r
  •  5
  • DeduciveR  · 技术社区  · 7 年前

    df <- structure(list(gender = c("female", "male", NA, NA, "male", "male", 
    "male"), Division = c("South Atlantic", "East North Central", 
    "Pacific", "East North Central", "South Atlantic", "South Atlantic", 
    "Pacific"), Median = c(57036.6262, 39917, 94060.208, 89822.1538, 
    107683.9118, 56149.3217, 46237.265), first_name = c("Marilyn", 
    "Jeffery", "Yashvir", "Deyou", "John", "Jose", "Daniel")), row.names = c(NA, 
    -7L), class = c("tbl_df", "tbl", "data.frame"))
    

    我需要进行分析 NA 中的值 gender 变量。其他列太少,而且没有已知的预测值,因此不可能对值进行插补。

    female male 找到失踪的案子。

    除了编写一些非常难看的代码来过滤不完整的情况外,可以一分为二地替换 与 女性的 男性的 在每一半中,我想知道是否有一种优雅的方法可以随机或按比例分配值到 不适用 是吗?

    3 回复  |  直到 7 年前
        1
  •  4
  •   jay.sf    7 年前

    我们可以利用 ifelse 和 is.na 确定 na 存在,然后使用 sample 随机选择 female male .

    df$gender <- ifelse(is.na(df$gender), sample(c("female", "male"), 1), df$gender)
    
        2
  •  4
  •   Santiago Capobianco    7 年前

    这个怎么样:

    > df <- structure(list(gender = c("female", "male", NA, NA, "male", "male", 
    +                                 "male"),
    +                      Division = c("South Atlantic", "East North Central", 
    +                                   "Pacific", "East North Central", "South Atlantic", "South Atlantic", 
    +                                   "Pacific"),
    +                      Median = c(57036.6262, 39917, 94060.208, 89822.1538,
    +                                 107683.9118, 56149.3217, 46237.265),
    +                      first_name = c("Marilyn", "Jeffery", "Yashvir", "Deyou", "John", "Jose", "Daniel")),
    +                 row.names = c(NA, -7L), class = c("tbl_df", "tbl", "data.frame"))
    > 
    > Gender <- rbinom(length(df$gender), 1, 0.52)
    > Gender <- factor(Gender, labels = c("female", "male"))
    > 
    > df$gender[is.na(df$gender)] <- as.character(Gender[is.na(df$gender)])
    > 
    > df$gender
    [1] "female" "male"   "female" "female" "male"   "male"   "male"  
    > 
    

    希望有帮助。

        3
  •  3
  •   BENY    7 年前

    只需分配

    df$gender[is.na(df$gender)]=sample(c("female", "male"), dim(df)[1], replace = TRUE)[is.na(df$gender)]