代码之家  ›  专栏  ›  技术社区  ›  Eliot Dixon tmfmnk

在dplyr::group_by中,获取多个分组变量之一的观测次数

  •  0
  • Eliot Dixon tmfmnk  · 技术社区  · 3 年前

    这很可能以前有人问过,但我很难阐明我的问题。

    在我的数据中,我有3个变量,LOCATION、TOPIC和RESPONSE。我想按位置计算TOPIC和RESPONSE的每个组合的分布。

    创建玩具数据并执行初始数据准备

    responses <- data.frame(LOCATION = c("LOC_A", "LOC_A", "LOC_A", "LOC_A", "LOC_A", 
                                         "LOC_A", "LOC_A", "LOC_A", 
                                         "LOC_B", "LOC_B", "LOC_B", "LOC_B", "LOC_B", 
                                         "LOC_C", "LOC_C", "LOC_C", "LOC_C", "LOC_C", "LOC_C", 
                                         "LOC_C", "LOC_C", "LOC_C", "LOC_C", "LOC_C", "LOC_C", 
                                         "LOC_C", "LOC_C", "LOC_C", "LOC_C", "LOC_C"),
                            TOPIC = c("Dogs", "Dogs", "Dogs", "Dogs", "Dogs", "Dogs", 
                                      "Lizards", "Lizards", "Lizards",
                                      "Lizards", "Lizards", "Lizards", "Lizards", "Lizards", 
                                       "Lizards", "Lizards", "Snakes", "Snakes", "Snakes", "Snakes", "Snakes", 
                                      "Snakes", "Dogs", "Snakes", "Dogs", "Snakes", "Dogs", 
                                      "Snakes", "Dogs", "Snakes"),
                            RESP = c("Agree", "Disagree", "Agree", "Disagree", "Agree", 
                                     "Disagree", "Agree", "Disagree", 
                                     "Agree", "Disagree", "Agree", "Disagree", "Neither", "Agree",
                                     "Neither", "Agree", "Neither", "Agree", "Neither", 
                                     "Agree", "Neither", "Agree", "Agree", "Neither", 
                                     "Agree", "Neither", "Agree", "Disagree", "Disagree",
                                     "Neither"))
    
    # Obtain counts for each combination of levels
    distribution <- responses %>% 
      table() %>% 
      as.data.frame() %>% 
      # Make it more readable
      dplyr::arrange(LOCATION, TOPIC, RESP) 
    
    

    下面是一个使用循环创建所需输出的示例解决方案:

    # ugly loop solution :(
    # Initialize output container
    out <- list()
    # Iterate over each location
    for(loc in unique(distribution$LOCATION)){
      # Subset distribution for this location
      thisDist <- dplyr::filter(distribution, LOCATION == loc)
      
      # Calculate percent of each response for this location
      thisDist$percent <- thisDist$Freq/sum(thisDist$Freq)
      
      # Store distribution df with percent column
      out[[loc]] <- thisDist
    
    }
    
    # combine output into single df
    out <- do.call("rbind", out)
    

    我想要的是一个简洁明了的解决方案。下面是一些伪代码,它描述了我想象中的解决方案。

    # Imaginary tidyverse solution :)
    out <- distribution %>% 
      group_by(LOCATION, TOPIC, RESP) %>% 
      summarise(#percent = Freq/(sum(<all-Freq-values-for-this-group's-LOCATION-value>))
                )
    

    我在这里要做的是获得当前组的LOCATION值的所有Freq值的总和。有没有一种很好的方法可以在一个小组中做到这一点?

    谢谢你的阅读,我希望这不是完全无法理解的。

    2 回复  |  直到 3 年前
        1
  •  1
  •   r2evans    3 年前

    这就是你要找的吗?

    distribution %>%
      mutate(percent = Freq/sum(Freq), .by = LOCATION)
    #    LOCATION   TOPIC     RESP Freq    percent
    # 1     LOC_A    Dogs    Agree    3 0.37500000
    # 2     LOC_A    Dogs Disagree    3 0.37500000
    # 3     LOC_A    Dogs  Neither    0 0.00000000
    # 4     LOC_A Lizards    Agree    1 0.12500000
    # 5     LOC_A Lizards Disagree    1 0.12500000
    # 6     LOC_A Lizards  Neither    0 0.00000000
    # 7     LOC_A  Snakes    Agree    0 0.00000000
    # 8     LOC_A  Snakes Disagree    0 0.00000000
    # 9     LOC_A  Snakes  Neither    0 0.00000000
    # 10    LOC_B    Dogs    Agree    0 0.00000000
    # 11    LOC_B    Dogs Disagree    0 0.00000000
    # 12    LOC_B    Dogs  Neither    0 0.00000000
    # 13    LOC_B Lizards    Agree    2 0.40000000
    # 14    LOC_B Lizards Disagree    2 0.40000000
    # 15    LOC_B Lizards  Neither    1 0.20000000
    # 16    LOC_B  Snakes    Agree    0 0.00000000
    # 17    LOC_B  Snakes Disagree    0 0.00000000
    # 18    LOC_B  Snakes  Neither    0 0.00000000
    # 19    LOC_C    Dogs    Agree    3 0.17647059
    # 20    LOC_C    Dogs Disagree    1 0.05882353
    # 21    LOC_C    Dogs  Neither    0 0.00000000
    # 22    LOC_C Lizards    Agree    2 0.11764706
    # 23    LOC_C Lizards Disagree    0 0.00000000
    # 24    LOC_C Lizards  Neither    1 0.05882353
    # 25    LOC_C  Snakes    Agree    3 0.17647059
    # 26    LOC_C  Snakes Disagree    1 0.05882353
    # 27    LOC_C  Snakes  Neither    6 0.35294118
    

    如果你有 dplyr 1.1以上,然后使用

    distribution %>%
      group_by(LOCATION) %>%
      mutate(percent = Freq/sum(Freq))
    
        2
  •  1
  •   William Wong    3 年前

    钥匙不用 summarise 但是 mutate 。

    out <- distribution %>% 
      ungroup() %>% 
      group_by(LOCATION) %>% 
      mutate(percent = Freq/ sum(Freq))
    
    推荐文章