代码之家  ›  专栏  ›  技术社区  ›  jdc0589

Linux shell脚本,用于计算文本文件中字符序列的出现次数?

  •  1
  • jdc0589  · 技术社区  · 16 年前

    例子: 如果我正在搜索“thisIsTheSequence”,则以下文件将有3个匹配项:

    asdasdthisIsTheSequence
    asdasdasthisIsT
    heSequenceasdasdthisIsTheSequ
    encesadasdasda
    

    谢谢你的帮助。

    4 回复  |  直到 16 年前
        1
  •  2
  •   ghostdog74    16 年前

    只需一个awk脚本即可,因为您将处理一个巨大的文件。使用多个管道可以降低速度。

    #!/bin/bash
    awk 'BEGIN{
     search="thisIsTheSequence"
     total=0
    }
    NR%10==0{
      c=gsub(search,"",s)
      total+=c  
    }
    NR{ s=s $0 }
    END{ 
     c=gsub(search,"",s)
     print "total count: "total+c
    }' file
    

    输出

    $ more file
    asdasdthisIsTheSequence
    asdasdasthisIsT
    heSequenceasdasdthisIsTheSequ
    encesadasdasdaasdasdthisIsTheSequence
    asdasdasthisIsT
    heSequenceasdasdthisIsTheSequ
    encesadasdasda
    asdasdthisIsTheSequence
    asdasdasthisIsT
    heSequenceasdasdthisIsTheSequ
    encesadasdasda
    
    $ ./shell.sh
    total count: 9
    
        2
  •  7
  •   bdonlan    16 年前

    echo $((`tr -d "\n" < file | sed 's/thisIsTheSequence/\n/g' | wc -l` - 1))
    

    在shell内核之外可能有更有效的方法使用实用程序——特别是如果您可以将文件放入内存中。

        3
  •  0
  •   Artelius    16 年前

    在你的序列中会有不止一条换行吗?

    如果没有,一种解决方案是将序列一分为二并搜索一半(例如,搜索“thisIsTh”和“eSequence”),然后返回您找到的事件并“仔细查看”,即去掉该区域中的换行符并检查是否匹配。

    基本上,这是一种快速“过滤”数据以找到有趣的东西。

        4
  •  -1
  •   Preet Sangha    16 年前

    使用类似于:

    head -n LL filename | tail -n YY | grep text | wc -l
    

    其中LL是序列的最后一行,YY是序列中的行数(即LL-第一行)