代码之家  ›  专栏  ›  技术社区  ›  Samizdis

从许多文本文件中快速删除前n行

  •  3
  • Samizdis  · 技术社区  · 16 年前

    我需要通过删除输入文件的前两行来创建输出文本文件。

    目前我正在使用 sed“1,2d”input.txt>output.txt

    我需要对数千个文件执行此操作,因此使用python:

    import os
    for filename in somelist:
      os.system('sed "1,2d" %s-in.txt > %s-out.txt'%(filename,filename))
    

    但这很慢。

    我需要保留原始文件,所以我不能把它放在适当的地方。

    有什么方法可以更快地做到这一点吗?使用的不是SED?也许使用其他脚本语言而不是Python?写一个短的C程序是值得的,还是文件写入磁盘访问可能是瓶颈?

    3 回复  |  直到 16 年前
        1
  •  9
  •   Cascabel    16 年前

    使用 tail . 怀疑任何事情都可能更快:

    tail -n +3 input.txt > output.txt
    

    把它包在你选择的循环中。但我真的怀疑SED慢了整整一吨——正如您所说,磁盘I/O通常是最终的瓶颈。

        2
  •  4
  •   nosklo    16 年前

    我认为这将比启动SED更快:

    import os
    import shutil
    
    path = '/some/path/to/files/'
    for filename in os.listdir(path):
        basename, ext = os.path.splitext(filename)
        fullname = os.path.join(path, filename)
        newname = os.path.join(path, basename + '-out' + ext)
        with open(fullname) as read:
            #skip first two lines
            for n in xrange(2):
                read.readline()
            # hand the rest to shutil.copyfileobj
            with open(newname, 'w') as write:
                shutil.copyfileobj(read, write)
    
        3
  •  3
  •   ghostdog74    16 年前
    for file in *.ext
    do
        sed -i.bak -n '3,$p' $file 
    done
    

    或者只是

    sed -i.bak -n '3,$p' *.ext