代码之家  ›  专栏  ›  技术社区  ›  Mahsa Hassankashi

如何通过sys读取Python中的两个不同文件。标准DIN

  •  1
  • Mahsa Hassankashi  · 技术社区  · 8 年前

    我想从sys读取两个不同的文件。stdin,我可以读写文件,但和第一个和第二个文件并没有分离。

    当我在cmd win 10和python 3.6上运行以下代码时:

    D:\digit>cat s.csv s2.csv
    

    结果是:

    1
    2
    3
    4
    5
    1
    2
    3
    4
    5
    6
    7
    

    我可以打印这两个文件。

    我的python代码是:

    import sys 
    import numpy as np
    
    train=[]
    test=[]
    
    #Assume below code is function 1 which just and must read s.csv
    reader = sys.stdin.readlines()
    for row in reader:          
        train.append(int(row[0]))
    train = np.array(train)
    
    print(train)
    
    #I need some thing here to make separation
    #sys.stdin.close()
    #sys.stdin = sys.__stdin__ 
    #sys.stdout.flush() 
    
    #Assume below code is function 2 which just and must read s2.csv
    reader = sys.stdin.readlines()
    for row in reader:          
        test.append(int(row[0]))
    test = np.array(test)
    
    print(test)
    

    我在cmd提示符下运行以下命令:

    D:\digit>cat s.csv s2.csv | python pytest.py
    

    结果是:

    [1 2 3 4 5 1 2 3 4 5 6 7]
    []
    

    我需要重置系统吗。下一个文件的stdin? 我使用了以下选项,但没有一个是答案:

    sys.stdin.close()
    sys.stdin = sys.__stdin__ 
    sys.stdout.flush() 
    

    非常感谢您的帮助。

    2 回复  |  直到 8 年前
        1
  •  2
  •   Edwin van Mierlo    8 年前

    让我试着解释一下。

    d:\digit>cat s.csv s2.csv
    

    只有1个输出,而不是2个。它的功能是“流式传输”内容 file1 到 stdout 然后将 file2 到 标准装置 , 没有任何停顿或分隔符!!

    因此只有1个“流”输出,然后使用|重定向到pyton脚本:

    | pytest.py
    

    所以 pytest.py 将接收1个“流”输入,它不知道更好或更多。

    如果要通过单独处理文件 pytest。py公司 ,可以执行以下操作

    D:\digit>cat s.csv | python pytest.py # process the first file
    D:\digit>cat s2.csv | python pytest.py # process the second file
    

    或在一个班轮上:

    D:\digit>cat s.csv | python pytest.py && cat s2.csv | python pytest.py
    

    只要记住 pytest。py公司 实际上正在运行 两次 . 因此,您需要为此调整python脚本。

    但是当您编辑python脚本时。。。

    您应该做什么: 如果要在 pytest。py公司 ,然后应该编写一些代码来读取python脚本中的两个文件。如果是csv结构化数据,请查看 csv module for reading and writing csv files

    [根据评论编辑:]

    我可以通过pandas“pd.read\u csv”读取多个文件,但我的 问题是我如何通过sys来完成它。stdin?

    你真的应该质疑为什么你如此专注于使用 stdin . 从python脚本中读取它可能会更加有效。

    如果必须使用 标准DIN 然后,您可以部署各种各样的、但位于python外部的、页眉、页脚和分隔符。一旦定义并能够这样做,就可以根据接收到的页眉/页脚/分隔符更改python中的代码以执行各种函数 标准DIN .

    这一切听起来有点复杂,容易出错。我强烈建议您重新考虑使用stdin作为脚本的输入。或者,请更新您的问题,说明您面临的技术要求和限制,这些限制了您使用stdin。

    [根据评论编辑:]

    我想加载这些文件我是Hadoop生态系统,我正在使用Hadoop 为此流媒体

    不知何故,您需要向python脚本发出“信号”,表明它正在处理一个新文件和新信息。

    假设您有2个文件,第一行需要是某种类型的“头”,指示文件,以及在收到新的“头”之前,需要对其余数据执行哪个函数。

    因此,假设您的“train”数据以行为前缀 @is_train@ 并且您的“测试”数据以行为前缀 @is_test@

    如何在您的环境中做到这一点,不属于此问题的范围

    现在,重定向到stdin将在数据之前发送这两个标头。您可以让python检查这些,例如:

    import sys 
    import numpy as np
    
    train=[]
    test=[]
    
    is_train = False
    is_test = False
    
    while True:
        line = sys.stdin.readline()
        if '@stop@' in line:
            break
        if '@is_train@' in line:
            is_train = True
            is_test = False
            continue
        if '@is_test@' in line:
            is_train = False
            is_test = True
            continue
        #if this is csv data, you might want to split on ,
        line = line.split(',')
        if is_train:
            train.append(int(line[0]))
        if is_test:
            test.append(int(line[0]))
    
    test = np.array(test)
    train = np.array(train)
    
    print(train)
    print(test)
    

    正如您在代码中看到的,在本例中,您还需要一个“页脚”来确定数据何时结束 @stop@ 已选择。

    发送页眉/页脚的一种方式可以是:

    D:\digit>cat is_train.txt s.csv is_test.txt s2.csv stop.txt | python pytest.py
    

    而这三个额外的文件只包含适当的页眉或页脚

        2
  •  2
  •   Mahsa Hassankashi    8 年前

    另一种解决方案是:

    import sys
    
    train=[]
    
    args = sys.stdin.readlines()[0].replace("\"", "").split()
    
    for arg in args:
        arg=arg.strip()
        with open(arg, "r") as f:
            train=[]
            for line in f:
                train.append(int(line))   
            print(train)    
    

    s、 txt是:

    1
    2
    3
    

    s2.txt是:

    7
    8
    9
    
    D:\digit>echo s.txt s2.txt | python argpy.py
    [1, 2, 3]
    [7, 8, 9]
    

    关键在于两点:

    1. 使用echo代替cat以防止级联 更多学习链接: Difference between 'cat < file.txt' and 'echo < file.txt'

    2. 通过拆分每个文件并存储在args中,尝试为每个新文件读入for循环。 How to run code with sys.stdin as input on multiple text files

    快乐bc我做到了:)