代码之家  ›  专栏  ›  技术社区  ›  Zaid

如何从数据框中的每个子类别中提取“最佳”行?[副本]

  •  1
  • Zaid  · 技术社区  · 9 年前

    在下面的示例中,由 order 函数用于按以下方式对每组中的条目进行排序:

    set.seed(123)
    
    ex.df <- data.frame(
      group = sample(LETTERS[1:4],20,replace=TRUE),
      score1 = sample(1:10),
      score2 = sample(1:10)
    )
    
    sortedOrderings <- by(ex.df, ex.df$group, function(df) order(df$score1 + df$score2) )
    
    bestIndices <- lapply(sortedOrderings, FUN= function(lst) lst[1] )
    

    查看数据帧的索引,该索引由 by ex.df ex.df公司 这不是最聪明的想法:

    print(sortedOrderings)
    
    ex.df$group: A
    [1] 2 3 4 1
    --------------------------------------------------------------- 
    ex.df$group: B
    [1] 5 3 2 4 1
    --------------------------------------------------------------- 
    ex.df$group: C
    [1] 2 1 3 4
    --------------------------------------------------------------- 
    ex.df$group: D
    [1] 3 7 4 6 1 2 5
    
    > print(ex.df[bestIndices,])
        group score1 score2
    2       D      7      9
    5       D      4      1
    2.1     D      7      9
    3       B      6      6
    

    有没有办法从小组中选出“最佳”一行 ,或至少有索引参考 ?

    2 回复  |  直到 9 年前
        1
  •  1
  •   Matt Summersgill    9 年前

    使用 data.table 要对第一行的索引执行自连接,其中总分等于分组的最大分数:

    set.seed(123)
    
    ex.df <- data.frame(
      group = sample(LETTERS[1:4],20,replace=TRUE),
      score1 = sample(1:10),
      score2 = sample(1:10)
    )
    
    library(data.table)
    setDT(ex.df)
    
    ex.df[ex.df[,.I[(score1 + score2) == max(score1 + score2)][1],by = .(group)]$V1][order(group)]
    

       group score1 score2
    1:     A      8      3
    2:     B      9     10
    3:     C     10      8
    4:     D      9     10
    
        2
  •  1
  •   tbradley    9 年前

    您可以使用 dplyr 程序包和 rank

    ex.df %>%
      mutate(total_score = score1 + score2) %>%
      group_by(group) %>%
      mutate(rank = rank(total_score)) %>%
      filter(rank == max(rank)) %>%
      select(-c(rank)) %>%
      arrange(group)
    

    给你这个:

    # A tibble: 4 x 4
    # Groups:   group [4]
       group score1 score2 total_score
      <fctr>  <int>  <int>       <int>
    1      A      8      3          11
    2      B      9     10          19
    3      C     10      8          18
    4      D      9     10          19