代码之家  ›  专栏  ›  技术社区  ›  apeescape

将不均匀分层列表转换为数据帧

  •  3
  • apeescape  · 技术社区  · 16 年前

    我认为还没有人问过这个问题,但有没有办法将多层次、结构不均匀的列表信息组合成一个“长”格式的数据框架?

    明确地:

    library(XML)
    library(plyr)
    xml.inning <- "http://gd2.mlb.com/components/game/mlb/year_2009/month_05/day_02/gid_2009_05_02_chamlb_texmlb_1/inning/inning_5.xml"
    xml.parse <- xmlInternalTreeParse(xml.inning)
    xml.list <- xmlToList(xml.parse)
    ## $top$atbat
    ## $top$atbat$pitch
    ##             des              id            type               x               y 
    ##          "Ball"           "310"             "B"         "70.39"        "125.20" 
    

    其中,结构如下所示:

    > llply(xml.list, function(x) llply(x, function(x) table(names(x))))
    $top
    $top$atbat
    .attrs  pitch 
         1      4 
    $top$atbat
    .attrs  pitch 
         1      4 
    $top$atbat
    .attrs  pitch 
         1      5 
    $bottom
    $bottom$action
         b    des  event      o  pitch player      s 
         1      1      1      1      1      1      1 
    $bottom$atbat
    .attrs  pitch 
         1      5 
    $bottom$atbat
    .attrs  pitch 
         1      5 
    $bottom$atbat
    .attrs  pitch runner 
         1      5      1 
    $bottom$atbat
    .attrs  pitch runner 
         1      7      1 
    $.attrs
    $.attrs$num
    character(0)
    $.attrs$away_team
    character(0)
    $.attrs$
    

    我想要的是一个数据帧,来自 类别,以及适当的( 顶部 , 阿特巴特 , 底部 ).因此,我需要忽略那些与数据不符的级别。由于列数不同而导致的帧。大概是这样的:

       first second third    des     x
    1    top  atbat pitch   Ball 70.29
    2    top  atbat pitch Strike 69.24
    3 bottom  atbat pitch    Out 67.22
    

    有没有优雅的方法?谢谢

    2 回复  |  直到 16 年前
        1
  •  5
  •   Joshua Ulrich    11 年前

    我不知道什么是优雅,但这很管用。那些更熟悉plyr的人可能会提供一个更通用的解决方案。

    cleanFun <- function(x) {
       a <- x[["atbat"]]
       b <- do.call(rbind,a[names(a)=="pitch"])
       c <- as.data.frame(b)
    }
    ldply(xml.list[c("top","bottom")], cleanFun)[,1:5]
         .id             des  id type      x
    1    top            Ball 310    B  70.39
    2    top   Called Strike 311    S 118.45
    3    top   Called Strike 312    S  86.70
    4    top In play, out(s) 313    X  79.83
    5 bottom            Ball 335    B  15.45
    6 bottom   Called Strike 336    S  77.25
    7 bottom Swinging Strike 337    S  99.57
    8 bottom            Ball 338    B 106.44
    9 bottom In play, out(s) 339    X 134.76
    
        2
  •  1
  •   apeescape    16 年前

    这个 .id 功能为 ldply() 这很好,但一旦你做了另一件事,它们似乎就重叠了 ldply() .

    下面是使用 rbind.fill() :

    aho <- ldply(llply(xml.list[[1]], function(x) ldply(x, function(x) rbind.fill(data.frame(t(x))))))
    > aho[1:5,1:4]
         .id                                                       des   id type
    1  pitch                                                      Ball  310    B
    2  pitch                                             Called Strike  311    S
    3  pitch                                             Called Strike  312    S
    4  pitch                                           In play, out(s)  313    X
    5 .attrs Alexei Ramirez lines out to second baseman Ian Kinsler.   <NA> <NA>
    

    这个 身份证件 第二次 ldply() 因为我们已经有了一个 身份证件 .我们可以通过命名第一个 身份证件 作为一个不同的名字,但它似乎并不连贯。

    aho2 <- ldply(llply(xml.list[[1]], function(x) {
      out <- ldply(x, function(x) rbind.fill(data.frame(t(x))))
      names(out)[1] <- ".id2"
      out
    }))
    > aho2[1:5,1:4]
        .id   .id2                                                       des   id
    1 atbat  pitch                                                      Ball  310
    2 atbat  pitch                                             Called Strike  311
    3 atbat  pitch                                             Called Strike  312
    4 atbat  pitch                                           In play, out(s)  313
    5 atbat .attrs Alexei Ramirez lines out to second baseman Ian Kinsler.   <NA>