代码之家  ›  专栏  ›  技术社区  ›  quamrana Ryuzaki L

如何在python中拆分一个巨大的文本文件

  •  21
  • quamrana Ryuzaki L  · 技术社区  · 17 年前

    我有一个巨大的文本文件(~1GB),遗憾的是,我使用的文本编辑器无法读取如此大的文件。然而,如果我能把它分成两到三个部分,我就可以了,因此,作为一个练习,我想用python编写一个程序来完成它。

    我想我想让程序做的是找到一个文件的大小,把这个数字分成几个部分,对于每个部分,以块的形式读到那个点,然后写到一个 .nnn输出文件,然后读取到下一个换行符并写入,然后关闭输出文件,等等。显然,最后一个输出文件只是复制到输入文件的末尾。

    我将首先编写这个代码测试,所以没有必要给我一个完整的答案,除非它是一行;-)

    14 回复  |  直到 17 年前
        1
  •  41
  •   James    16 年前

    linux有一个split命令

        2
  •  16
  •   Kamil Kisiel    17 年前

    退房 os.stat() file.readlines([sizehint])

        3
  •  9
  •   Alex L    14 年前

    作为替代方法,使用日志库:

    >>> import logging.handlers
    >>> log = logging.getLogger()
    >>> fh = logging.handlers.RotatingFileHandler("D://filename.txt", 
         maxBytes=2**20*100, backupCount=100) 
    # 100 MB each, up to a maximum of 100 files
    >>> log.addHandler(fh)
    >>> log.setLevel(logging.INFO)
    >>> f = open("D://biglog.txt")
    >>> while True:
    ...     log.info(f.readline().strip())
    

    您的文件将显示如下:

    filename.txt(文件末尾)
    filename.txt.1
    filename.txt.2

    filename.txt.10(文件开头)

    这是一种快速简便的方法,可以使一个巨大的日志文件与您的 RotatingFileHandler 实施

        4
  •  9
  •   Ram    8 年前
        5
  •  6
  •   Joe Koberg    16 年前

    别忘了 seek() 和 mmap() 用于随机访问文件。

    def getSomeChunk(filename, start, len):
        fobj = open(filename, 'r+b')
        m = mmap.mmap(fobj.fileno(), 0)
        return m[start:start+len]
    
        6
  •  6
  •   Ryan Ginstrom    14 年前

    这种生成器方法是一种(缓慢的)方式,可以在不破坏内存的情况下获得一段线。

    import itertools
    
    def slicefile(filename, start, end):
        lines = open(filename)
        return itertools.islice(lines, start, end)
    
    out = open("/blah.txt", "w")
    for line in slicefile("/python27/readme.txt", 10, 15):
        out.write(line)
    
        7
  •  4
  •   Community Mohan Dere    9 年前

    虽然 Ryan Ginstrom's answer itertools.islice 通过依次迭代打开的文件描述符:

    def splitfile(infilepath, chunksize):
        fname, ext = infilepath.rsplit('.',1)
        i = 0
        written = False
        with open(infilepath) as infile:
            while True:
                outfilepath = "{}{}.{}".format(fname, i, ext)
                with open(outfilepath, 'w') as outfile:
                    for line in (infile.readline() for _ in range(chunksize)):
                        outfile.write(line)
                    written = bool(line)
                if not written:
                    break
                i += 1
    
        8
  •  2
  •   Svante    17 年前

    你可以用 wc 和 split bash :

    split -dl$((`wc -l 'filename'|sed 's/ .*$//'` / 3 + 1)) filename filename-chunk.
    

    生成相同行数的3个部分(当然最后一个部分有舍入错误),命名为 filename-chunk.00 到 filename-chunk.02 .

        9
  •  2
  •   quamrana Ryuzaki L    17 年前

    我已经写了这个程序,看起来效果不错。谢谢卡米尔·基希尔让我开始学习。
    (请注意,FileSizeParts()是此处未显示的函数)
    稍后我可能会花时间做一个二进制读取的版本,看看它是否更快。

    def Split(inputFile,numParts,outputName):
        fileSize=os.stat(inputFile).st_size
        parts=FileSizeParts(fileSize,numParts)
        openInputFile = open(inputFile, 'r')
        outPart=1
        for part in parts:
            if openInputFile.tell()<fileSize:
                fullOutputName=outputName+os.extsep+str(outPart)
                outPart+=1
                openOutputFile=open(fullOutputName,'w')
                openOutputFile.writelines(openInputFile.readlines(part))
                openOutputFile.close()
        openInputFile.close()
        return outPart-1
    
        10
  •  2
  •   Mudit Verma    10 年前

    用法-split.py文件名splitsizeinkb

    import os
    import sys
    
    def getfilesize(filename):
       with open(filename,"rb") as fr:
           fr.seek(0,2) # move to end of the file
           size=fr.tell()
           print("getfilesize: size: %s" % size)
           return fr.tell()
    
    def splitfile(filename, splitsize):
       # Open original file in read only mode
       if not os.path.isfile(filename):
           print("No such file as: \"%s\"" % filename)
           return
    
       filesize=getfilesize(filename)
       with open(filename,"rb") as fr:
        counter=1
        orginalfilename = filename.split(".")
        readlimit = 5000 #read 5kb at a time
        n_splits = filesize//splitsize
        print("splitfile: No of splits required: %s" % str(n_splits))
        for i in range(n_splits+1):
            chunks_count = int(splitsize)//int(readlimit)
            data_5kb = fr.read(readlimit) # read
            # Create split files
            print("chunks_count: %d" % chunks_count)
            with open(orginalfilename[0]+"_{id}.".format(id=str(counter))+orginalfilename[1],"ab") as fw:
                fw.seek(0) 
                fw.truncate()# truncate original if present
                while data_5kb:                
                    fw.write(data_5kb)
                    if chunks_count:
                        chunks_count-=1
                        data_5kb = fr.read(readlimit)
                    else: break            
            counter+=1 
    
    if __name__ == "__main__":
       if len(sys.argv) < 3: print("Filename or splitsize not provided: Usage:     filesplit.py filename splitsizeinkb ")
       else:
           filesize = int(sys.argv[2]) * 1000 #make into kb
           filename = sys.argv[1]
           splitfile(filename, filesize)
    
        11
  •  2
  •   radtek    6 年前

    下面是一个python脚本,可用于使用 subprocess :

    """
    Splits the file into the same directory and
    deletes the original file
    """
    
    import subprocess
    import sys
    import os
    
    SPLIT_FILE_CHUNK_SIZE = '5000'
    SPLIT_PREFIX_LENGTH = '2'  # subprocess expects a string, i.e. 2 = aa, ab, ac etc..
    
    if __name__ == "__main__":
    
        file_path = sys.argv[1]
        # i.e. split -a 2 -l 5000 t/some_file.txt ~/tmp/t/
        subprocess.call(["split", "-a", SPLIT_PREFIX_LENGTH, "-l", SPLIT_FILE_CHUNK_SIZE, file_path,
                         os.path.dirname(file_path) + '/'])
    
        # Remove the original file once done splitting
        try:
            os.remove(file_path)
        except OSError:
            pass
    

    您可以在外部称之为:

    import os
    fs_result = os.system("python file_splitter.py {}".format(local_file_path))
    

    子流程 创建一个内存占用与进程大小相同的fork,如果进程内存已经很重,则在运行时将其加倍。同样的事情 os.system .

    下面是另一种纯python的方法,虽然我没有在大型文件上测试过,但速度会慢一些,但内存会更精简:

    CHUNK_SIZE = 5000
    
    def yield_csv_rows(reader, chunk_size):
        """
        Opens file to ingest, reads each line to return list of rows
        Expects the header is already removed
        Replacement for ingest_csv
        :param reader: dictReader
        :param chunk_size: int, chunk size
        """
        chunk = []
        for i, row in enumerate(reader):
            if i % chunk_size == 0 and i > 0:
                yield chunk
                del chunk[:]
            chunk.append(row)
        yield chunk
    
    with open(local_file_path, 'rb') as f:
        f.readline().strip().replace('"', '')
        reader = unicodecsv.DictReader(f, fieldnames=header.split(','), delimiter=',', quotechar='"')
        chunks = yield_csv_rows(reader, CHUNK_SIZE)
        for chunk in chunks:
            if not chunk:
                break
            # Do something with your chunk here
    

    readlines() :

    """
    Simple example using readlines()
    where the 'file' is generated via:
    seq 10000 > file
    """
    CHUNK_SIZE = 5
    
    
    def yield_rows(reader, chunk_size):
        """
        Yield row chunks
        """
        chunk = []
        for i, row in enumerate(reader):
            if i % chunk_size == 0 and i > 0:
                yield chunk
                del chunk[:]
            chunk.append(row)
        yield chunk
    
    
    def batch_operation(data):
        for item in data:
            print(item)
    
    
    with open('file', 'r') as f:
        chunks = yield_rows(f.readlines(), CHUNK_SIZE)
        for _chunk in chunks:
            batch_operation(_chunk)
    

    readlines示例演示如何对数据进行分块,以便将分块传递给需要分块的函数。不幸的是,readlines在内存中打开了整个文件,为了提高性能,最好使用reader示例。虽然如果您可以轻松地将所需内容放入内存,并需要分块处理,这就足够了。

        12
  •  1
  •   Ryan    12 年前

    import os
    
    fil = "inputfile"
    outfil = "outputfile"
    
    f = open(fil,'r')
    
    numbits = 1000000000
    
    for i in range(0,os.stat(fil).st_size/numbits+1):
        o = open(outfil+str(i),'w')
        segment = f.readlines(numbits)
        for c in range(0,len(segment)):
            o.write(segment[c]+"\n")
        o.close()
    
        13
  •  1
  •   Ajith    6 年前

    您可以实现将任何文件拆分为块,如下所示,块大小为500000字节(500kb),内容可以是任何文件:

    for idx,val in enumerate(get_chunk(content, CHUNK_SIZE)):
        data=val
        index=idx
    
    def get_chunk(content,size):
            for i in range(0,len(content),size):
                yield content[i:i+size]
    
        14
  •  0
  •   Ron Smith    12 年前

    我需要拆分csv文件以导入Dynamics CRM,因为导入的文件大小限制为8MB,并且我们收到的文件要大得多。此程序允许用户输入文件名和LinesPerFile,然后将指定的文件拆分为请求的行数。我真不敢相信它的速度有多快!

    # user input FileNames and LinesPerFile
    FileCount = 1
    FileNames = []
    while True:
        FileName = raw_input('File Name ' + str(FileCount) + ' (enter "Done" after last File):')
        FileCount = FileCount + 1
        if FileName == 'Done':
            break
        else:
            FileNames.append(FileName)
    LinesPerFile = raw_input('Lines Per File:')
    LinesPerFile = int(LinesPerFile)
    
    for FileName in FileNames:
        File = open(FileName)
    
        # get Header row
        for Line in File:
            Header = Line
            break
    
        FileCount = 0
        Linecount = 1
        for Line in File:
    
            #skip Header in File
            if Line == Header:
                continue
    
            #create NewFile with Header every [LinesPerFile] Lines
            if Linecount % LinesPerFile == 1:
                FileCount = FileCount + 1
                NewFileName = FileName[:FileName.find('.')] + '-Part' + str(FileCount) + FileName[FileName.find('.'):]
                NewFile = open(NewFileName,'w')
                NewFile.write(Header)
    
            NewFile.write(Line)
            Linecount = Linecount + 1
    
        NewFile.close()
    
        15
  •  -1
  •   Claudiu    17 年前

    或者,wc和split的python版本:

    lines = 0
    for l in open(filename): lines += 1
    

    然后编写一些代码,将第一行/3读入一个文件,将下一行/3读入另一个文件,以此类推。