代码之家  ›  专栏  ›  技术社区  ›  MarkAWard

Unix排序产生错误输出

  •  0
  • MarkAWard  · 技术社区  · 12 年前

    我正在尝试通过以下操作测试hadoop流作业的mapper和reducer函数:

        cat data.txt | python mapper.py | sort | python reducer.py
    

    但是来自映射器的排序输出不正确。

    he the  1
    i       1
    i dog   1
    i like  1
    i'm     1
    i'm rob 1
    i'm the 1
    i the   1 ### this should be after "i like 1" ###
    lazy    1
    

    我让其他人在他们的机器上测试过,他们用同样精确的映射器函数和命令行执行得到了正确的输出。所以我的Unix排序似乎出了问题。

    如果这有帮助:

    echo $TERM
    > vt100 
    

    对于尝试或设置不同的内容,我们将非常感激。谢谢

    1 回复  |  直到 12 年前
        1
  •  5
  •   Community Mohan Dere    9 年前

    你有你的答案 here 这是关于语言环境的。简而言之,您应该使用

    cat data.txt | python mapper.py | LC_COLLATE=C sort
    
    推荐文章