代码之家  ›  专栏  ›  技术社区  ›  Grey

标准xml解析器在Golang中的性能非常低

  •  5
  • Grey  · 技术社区  · 9 年前

    我有一个100Gb大小的xml文件,并在代码中使用SAX方法对其进行解析

    file, err := os.Open(filename)
    handle(err)
    defer file.Close()
    buffer := bufio.NewReaderSize(file, 1024*1024*256) // 33554432
    decoder := xml.NewDecoder(buffer)
    for {
            t, _ := decoder.Token()
            if t == nil {
                break
            }
            switch se := t.(type) {
            case xml.StartElement:
                if se.Name.Local == "House" {
                    house := House{}
                    err := decoder.DecodeElement(&house, &se)
                    handle(err)
                }
            }
        }
    

    但golang的工作速度很慢,这似乎取决于执行时间和磁盘使用情况。我的硬盘能够以大约100-120 mb/s的速度读取数据,但golang仅使用10-13 mb/s。 为了进行实验,我用c语言重写了这段代码:

    using (XmlReader reader = XmlReader.Create(filename)
                {
                    while (reader.Read())
                    {
                        switch (reader.NodeType)
                        {
                            case XmlNodeType.Element:
                                if (reader.Name == "House")
                                {
                                    //Code
                                }
                                break;
                        }
                    }
                }
    

    我得到了全硬盘加载,以100-110mb/s的速度读取数据。执行时间大约减少10倍。

    如何使用golang提高xml解析性能?

    2 回复  |  直到 6 年前
        1
  •  3
  •   Goodwine    6 年前

    这5件事可以帮助提高速度使用 encoding/xml 图书馆:
    (针对75k个条目的XMB进行测试,20MB,%s应用于前一个项目符号)

    1. 使生效 xml.Unmarshaller 在你所有的结构上
      • 很多代码
    2. 代替 d.DecodeElement(&foo, &token) 具有 foo.UnmarshallXML(d, &token)
      • 几乎100%安全
      • 节省10%的时间;allocs公司
    3. 使用 d.RawToken() d.Token()
      • 需要手动处理嵌套对象和命名空间
      • 节省10%的时间;20%allocs
    4. 如果使用,则使用 d.Skip()

    我 在我的特定用例中,代价是更多的代码、boileplate和可能更糟糕的角落案例处理,但我的输入相当一致,但这还不够。

    benchstat first.bench.txt parseraw.bench.txt 
    name          old time/op    new time/op    delta
    Unmarshal-16     1.06s ± 6%     0.66s ± 4%  -37.55%  (p=0.008 n=5+5)
    
    name          old alloc/op   new alloc/op   delta
    Unmarshal-16     461MB ± 0%     280MB ± 0%  -39.20%  (p=0.029 n=4+4)
    
    name          old allocs/op  new allocs/op  delta
    Unmarshal-16     8.42M ± 0%     5.03M ± 0%  -40.26%  (p=0.016 n=4+5)
    

    在我的实验中 lack of memoizing issue

        2
  •  2
  •   tsak    7 年前

    回答你的问题

    使用公共 xml.NewDecoder / decoder.Token ,我在本地看到50MB/s。通过使用 https://github.com/tamerh/xml-stream-parser

    为了测试我使用的 Posts.xml https://archive.org/details/stackexchange 归档torrent。

    package main
    
    import (
        "bufio"
        "fmt"
        "github.com/tamerh/xml-stream-parser"
        "os"
        "time"
    )
    
    func main() {
        // Using `Posts.xml` (68 GB) from https://archive.org/details/stackexchange (in the torrent)
        f, err := os.Open("Posts.xml")
        if err != nil {
            panic(err)
        }
        defer f.Close()
    
        br := bufio.NewReaderSize(f, 1024*1024)
        parser := xmlparser.NewXmlParser(br, "row")
    
        started := time.Now()
        var previous int64 = 0
    
        for x := range *parser.Stream() {
            elapsed := int64(time.Since(started).Seconds())
            if elapsed > previous {
                kBytesPerSecond := int64(parser.TotalReadSize) / elapsed / 1024
                fmt.Printf("\r%ds elapsed, read %d kB/s (last post.Id %s)", elapsed, kBytesPerSecond, x.Attrs["Id"])
                previous = elapsed
            }
        }
    }
    

    这将输出以下内容:

    ...s elapsed, read ... kB/s (last post.Id ...)
    

    唯一需要注意的是,这并没有为您提供方便的结构解组。

    https://github.com/golang/go/issues/21823 ,速度似乎是Golang中XML实现的一般问题,需要重写/重新思考标准库的这一部分。