代码之家  ›  专栏  ›  技术社区  ›  Ivo Flipse

如何将多维数组写入文本文件?

  •  93
  • Ivo Flipse  · 技术社区  · 15 年前

    在另一个问题中,如果我能提供我遇到问题的阵列,其他用户会提供一些帮助。然而,我甚至在一个基本的I/O任务上失败,比如将数组写入文件。

    这个数组由四个11 x 14的数组组成,所以我应该用一个漂亮的换行符来格式化它,以使其他人更容易读取文件。

    :所以我试过numpy.savetxt文件功能。奇怪的是,它给出了以下错误:

    TypeError: float argument required, not numpy.ndarray
    

    我假设这是因为函数不能处理多维数组?任何解决方案,因为我想他们在一个文件?

    8 回复  |  直到 15 年前
        1
  •  209
  •   Community Mohan Dere    5 年前

    如果您想将其写入磁盘,以便可以轻松地作为numpy数组读回,请查看 numpy.save . 酸洗也可以很好地工作,但对于大型阵列来说效率较低(而您的阵列不是,所以两者都很好)。

    numpy.savetxt .

    编辑: 所以,看起来 savetxt

    我才意识到 阻塞超过2维的数组。。。这可能是设计的,因为在文本文件中没有内在定义的方法来指示额外的维度。

    import numpy as np
    x = np.arange(20).reshape((4,5))
    np.savetxt('test.txt', x)
    

    而同样的事情会失败(有一个相当不具信息性的错误: TypeError: float argument required, not numpy.ndarray )对于三维阵列:

    import numpy as np
    x = np.arange(200).reshape((4,5,10))
    np.savetxt('test.txt', x)
    

    一种解决方法是将3D(或更大的)数组分解为2D切片。例如。

    x = np.arange(200).reshape((4,5,10))
    with open('test.txt', 'w') as outfile:
        for slice_2d in x:
            np.savetxt(outfile, slice_2d)
    

    numpy.loadtxt . 因此,我们可以更详细一点,并使用注释掉的行来区分切片。默认情况下, # (或以法律所指明者为准) comments

    import numpy as np
    
    # Generate some test data
    data = np.arange(200).reshape((4,5,10))
    
    # Write the array to disk
    with open('test.txt', 'w') as outfile:
        # I'm writing a header here just for the sake of readability
        # Any line starting with "#" will be ignored by numpy.loadtxt
        outfile.write('# Array shape: {0}\n'.format(data.shape))
        
        # Iterating through a ndimensional array produces slices along
        # the last axis. This is equivalent to data[i,:,:] in this case
        for data_slice in data:
    
            # The formatting string indicates that I'm writing out
            # the values in left-justified columns 7 characters in width
            # with 2 decimal places.  
            np.savetxt(outfile, data_slice, fmt='%-7.2f')
    
            # Writing out a break to indicate different slices...
            outfile.write('# New slice\n')
    

    这将产生:

    # Array shape: (4, 5, 10)
    0.00    1.00    2.00    3.00    4.00    5.00    6.00    7.00    8.00    9.00   
    10.00   11.00   12.00   13.00   14.00   15.00   16.00   17.00   18.00   19.00  
    20.00   21.00   22.00   23.00   24.00   25.00   26.00   27.00   28.00   29.00  
    30.00   31.00   32.00   33.00   34.00   35.00   36.00   37.00   38.00   39.00  
    40.00   41.00   42.00   43.00   44.00   45.00   46.00   47.00   48.00   49.00  
    # New slice
    50.00   51.00   52.00   53.00   54.00   55.00   56.00   57.00   58.00   59.00  
    60.00   61.00   62.00   63.00   64.00   65.00   66.00   67.00   68.00   69.00  
    70.00   71.00   72.00   73.00   74.00   75.00   76.00   77.00   78.00   79.00  
    80.00   81.00   82.00   83.00   84.00   85.00   86.00   87.00   88.00   89.00  
    90.00   91.00   92.00   93.00   94.00   95.00   96.00   97.00   98.00   99.00  
    # New slice
    100.00  101.00  102.00  103.00  104.00  105.00  106.00  107.00  108.00  109.00 
    110.00  111.00  112.00  113.00  114.00  115.00  116.00  117.00  118.00  119.00 
    120.00  121.00  122.00  123.00  124.00  125.00  126.00  127.00  128.00  129.00 
    130.00  131.00  132.00  133.00  134.00  135.00  136.00  137.00  138.00  139.00 
    140.00  141.00  142.00  143.00  144.00  145.00  146.00  147.00  148.00  149.00 
    # New slice
    150.00  151.00  152.00  153.00  154.00  155.00  156.00  157.00  158.00  159.00 
    160.00  161.00  162.00  163.00  164.00  165.00  166.00  167.00  168.00  169.00 
    170.00  171.00  172.00  173.00  174.00  175.00  176.00  177.00  178.00  179.00 
    180.00  181.00  182.00  183.00  184.00  185.00  186.00  187.00  188.00  189.00 
    190.00  191.00  192.00  193.00  194.00  195.00  196.00  197.00  198.00  199.00 
    # New slice
    

    numpy.loadtxt('test.txt').reshape((4,5,10)) . 举个例子(你可以用一行字来表达,我只是为了澄清一些事情):

    # Read the array from disk
    new_data = np.loadtxt('test.txt')
    
    # Note that this returned a 2D array!
    print new_data.shape
    
    # However, going back to 3D is easy if we know the 
    # original shape of the array
    new_data = new_data.reshape((4,5,10))
        
    # Just to check that they're the same...
    assert np.all(new_data == data)
    
        2
  •  33
  •   zyy    7 年前

    我不确定这是否符合您的要求,因为我认为您有兴趣使文件可读的人,但如果这不是一个主要的问题,只是 pickle 是的。

    import pickle
    
    my_data = {'a': [1, 2.0, 3, 4+6j],
               'b': ('string', u'Unicode string'),
               'c': None}
    output = open('data.pkl', 'wb')
    pickle.dump(my_data, output)
    output.close()
    

    重读:

    import pprint, pickle
    
    pkl_file = open('data.pkl', 'rb')
    
    data1 = pickle.load(pkl_file)
    pprint.pprint(data1)
    
    pkl_file.close()
    
        3
  •  11
  •   aseagram David Hagan    13 年前

    如果不需要可读的输出,另一种方法是将数组保存为MATLAB .mat .垫子 在很少的线路上是方便的。

    不需要知道数据的原始形状 文件,即在读入时无需重塑。而且,与使用 pickle ,一个 .垫子 文件可以由MATLAB读取,也可以由其他一些程序/语言读取。

    举个例子:

    import numpy as np
    import scipy.io
    
    # Some test data
    x = np.arange(200).reshape((4,5,10))
    
    # Specify the filename of the .mat file
    matfile = 'test_mat.mat'
    
    # Write the array to the mat file. For this to work, the array must be the value
    # corresponding to a key name of your choice in a dictionary
    scipy.io.savemat(matfile, mdict={'out': x}, oned_as='row')
    
    # For the above line, I specified the kwarg oned_as since python (2.7 with 
    # numpy 1.6.1) throws a FutureWarning.  Here, this isn't really necessary 
    # since oned_as is a kwarg for dealing with 1-D arrays.
    
    # Now load in the data from the .mat that was just saved
    matdata = scipy.io.loadmat(matfile)
    
    # And just to check if the data is the same:
    assert np.all(x == matdata['out'])
    

    如果忘记了数组在 .垫子 文件,您可以始终执行以下操作:

    print matdata.keys()
    

    所以是的,你的眼睛看不懂,但是只需要两行代码就可以写和读数据,我认为这是一个公平的权衡。

    看一下你的文件 scipy.io.savemat scipy.io.loadmat 以及本教程页面: scipy.io File IO Tutorial

        4
  •  9
  •   Lee    8 年前

    ndarray.tofile()

    e、 g.如果你的数组被调用 a :

    a.tofile('yourfile.txt',sep=" ",format="%s")
    

    编辑 (凯文J布莱克的评论) here ):

    从1.5.0版开始, np.tofile() 接受可选参数 newline='\n' 允许多行输出。 https://docs.scipy.org/doc/numpy-1.13.0/reference/generated/numpy.savetxt.html

        5
  •  4
  •   Ronny Brendel    15 年前

    有专门的图书馆可以做到这一点。(加上python的包装器)

    希望这有帮助

        6
  •  1
  •   jwueller    15 年前

        7
  •  0
  •   BennyD    12 年前

    我有一个简单的方法文件名.write()操作。它对我来说工作得很好,但是我要处理的数组有大约1500个数据元素。

    我基本上只需要for循环遍历文件,并以csv样式的输出逐行将其写入输出目标。

    import numpy as np
    
    trial = np.genfromtxt("/extension/file.txt", dtype = str, delimiter = ",")
    
    with open("/extension/file.txt", "w") as f:
        for x in xrange(len(trial[:,1])):
            for y in range(num_of_columns):
                if y < num_of_columns-2:
                    f.write(trial[x][y] + ",")
                elif y == num_of_columns-1:
                    f.write(trial[x][y])
            f.write("\n")
    

    if和elif语句用于在数据元素之间添加逗号。无论出于什么原因,当以nd数组的形式读入文件时,它们都会被剥离。我的目标是将文件输出为csv,所以这个方法有助于处理这个问题。

    希望这有帮助!

        8
  •  0
  •   Andrew Fan    7 年前

    泡菜最适合这种情况。假设你有一个叫 x_train

    import pickle
    
    ###Load into file
    with open("myfile.pkl","wb") as f:
        pickle.dump(x_train,f)
    
    ###Extract from file
    with open("myfile.pkl","rb") as f:
        x_temp = pickle.load(f)
    
        9
  •  0
  •   kenorb    5 年前

    对多维数组使用JSON模块,例如。

    import json
    with open(filename, 'w') as f:
       json.dump(myndarray.tolist(), f)
    
        10
  •  0
  •   kenorb    5 年前

    Write to a file with Python's print()

    import numpy as np
    import sys
    
    stdout_sys = sys.stdout
    np.set_printoptions(precision=8) # Sets number of digits of precision.
    np.set_printoptions(suppress=True) # Suppress scientific notations.
    np.set_printoptions(threshold=sys.maxsize) # Prints the whole arrays.
    with open('myfile.txt', 'w') as f:
        sys.stdout = f
        print(nparr)
        sys.stdout = stdout_sys
    

    使用 set_printoptions() to customize 对象的显示方式。

        11
  •  0
  •   Nico Schlömer David Maze    5 年前

    文件I/O通常是代码中的瓶颈。这就是为什么一定要知道ASCII/O总是比二进制I/O慢很多的原因 perfplot :

    enter image description here

    复制绘图的代码:

    import json
    import pickle
    
    import numpy as np
    import perfplot
    import scipy.io
    
    
    def numpy_save(data):
        np.save("test.dat", data)
    
    
    def numpy_savetxt(data):
        np.savetxt("test.txt", data)
    
    
    def numpy_savetxt_fmt(data):
        np.savetxt("test.txt", data, fmt="%-7.2f")
    
    
    def pickle_dump(data):
        with open("data.pkl", "wb") as f:
            pickle.dump(data, f)
    
    
    def scipy_savemat(data):
        scipy.io.savemat("test.dat", mdict={"out": data})
    
    
    def numpy_tofile(data):
        data.tofile("test.txt", sep=" ", format="%s")
    
    
    def json_dump(data):
        with open("test.json", "w") as f:
            json.dump(data.tolist(), f)
    
    
    perfplot.save(
        "out.png",
        setup=np.random.rand,
        n_range=[2 ** k for k in range(20)],
        kernels=[
            numpy_save,
            numpy_savetxt,
            numpy_savetxt_fmt,
            pickle_dump,
            scipy_savemat,
            numpy_tofile,
            json_dump,
        ],
        equality_check=None,
    )