代码之家  ›  专栏  ›  技术社区  ›  gabor

linux下如何区分二进制文件和文本文件

  •  12
  • gabor  · 技术社区  · 16 年前

    linux file 命令在识别文件类型方面做得非常好,并给出非常细粒度的结果。这个 diff 该工具能够区分二进制文件和文本文件,产生不同的输出。

    差异 将尝试基于文本的比较。

    为了澄清这个问题:我不在乎它是ASCII文本还是XML,只要它是文本就行。另外,我不想区分MP3和JPEG文件,因为它们都是二进制文件。

    8 回复  |  直到 16 年前
        1
  •  7
  •   David Schmitt    16 年前

    这个 diff manual

    diff确定文件是否为文本 字节依赖于系统,但是 通常有几千个。如果每 非null,diff将文件视为 成为文本;否则它认为 文件是二进制的。

        2
  •  12
  •   Tyler McHenry    16 年前

    file 仍然是你想要的命令。任何文本文件(根据其启发法)都将在输出中包含单词“text” ; 任何二进制文件都不包含“文本”一词。

    文件 文件 不将ASCII格式的PGP公钥块标识为“文本”,但可以(因为它仅由可打印字符组成,即使它不是人类可读的)。

        3
  •  6
  •   RichieHindle    16 年前

    一个快速而肮脏的方法是寻找 NUL 努尔 .

    更新

        4
  •  3
  •   Simone Margaritelli    16 年前

    你可以试试看

    strings yourfile
    

    命令并将结果大小与文件大小进行比较。。。我不完全确定,但如果它们是相同的文件实际上是一个文本文件。

        5
  •  3
  •   Robin A. Meade    5 年前

    这种方法服从于 grep 用于确定文件是二进制文件还是文本文件的命令:

    is_text_file() { grep -qIF '' "$1"; }
    

    • -q 安静的;如果发现任何匹配项,则立即以零状态退出
    • -I
    • -F 将模式解释为固定字符串,而不是正则表达式。

    使用的grep模式:

    • '' 空字符串。所有文件(空文件除外) 将匹配此模式。

    • file 司令部同意这一评估。)
    • 有一个可打印字符的文件,例如 a (对我来说很有意义。) 文件 司令部不同意这一评估(用GNU测试 文件
    • 这种方法只需要一个子进程来测试文件是文本还是二进制文件。

    试验

    # cd into a temp directory
    cd "$(mktemp -d)"
    
    # Create 3 corner-case test files
    touch empty_file       # An empty file
    echo -n a >one_byte_a  # A file containing just `a`
    echo a >one_line_a     # A file containing just `a` and a newline
    
    # Another test case: a 96KiB text file that ends with a NUL
    head -c 98303 /usr/share/dict/words > file_with_a_null_96KiB
    dd if=/dev/zero bs=1 count=1 >> file_with_a_null_96KiB
    
    # Last test case: a 96KiB text file plus a NUL added at the end
    head -c 98304 /usr/share/dict/words > file_with_a_null_96KiB_plus1
    dd if=/dev/zero bs=1 count=1 >> file_with_a_null_96KiB_plus1
    
    # Defer to grep to determine if a file is a text file
    is_text_file() { grep -qI '^' "$1"; }
    
    # Test harness
    do_test() {
      printf '%22s ... ' "$1"
      if is_text_file "$1"; then
        echo "is a text file"
      else
        echo "is a binary file"
      fi
    }
    
    # Test each of our test cases
    do_test empty_file
    do_test one_byte_a
    do_test one_line_a
    do_test file_with_a_null_96KiB
    do_test file_with_a_null_96KiB_plus1
    

                empty_file ... is a binary file
                one_byte_a ... is a text file
                one_line_a ... is a text file
    file_with_a_null_96KiB ... is a binary file
    file_with_a_null_96KiB_plus1 ... is a text file
    

    在我的机器上,似乎grep检查了一个文件的前96 KiB以查找 NUL 格雷普

    相关源代码: https://git.savannah.gnu.org/cgit/grep.git/tree/src/grep.c?h=v3.6#n1550

        6
  •  1
  •   Christoffer Hammarström    16 年前

    现在术语“文本文件”是模棱两可的,因为文本文件可以用ASCII、ISO-8859-*、UTF-8、UTF-16、UTF-32等编码。

    看到了吗 here

        7
  •  0
  •   yoshi    11 年前

        8
  •  -1
  •   Raghu    16 年前