我们不知道发生这种情况的请求id是哪个,文件中可能会发生什么变化。
先读,直到短语出现,将行保存在缓冲区中。一旦我们从该行解析请求id,我们就知道它是哪一个。然后处理累积的缓冲区,打印其他请求id的所有行,并且只打印超出“所选”请求的跳过距离的行。
然后继续将行收集到缓冲区中,直到再次找到短语。处理缓冲区:打印其他id的所有行,并且只打印那些距离所选id的两个短语足够远的行。
use warnings;
use strict;
use feature 'say';
my ($file, $skip_dist) = @ARGV;
die "Usage: $0 log-file [skip-distance]\n" if not $file;
$skip_dist //= 2; #/
my $trigger = qr{\[GET_REGION_INFO\]};
open my $fh, '<', $file or die "Can't open $file: $!";
my (@buf, $req_mark, $skip_idx, $next_req_cnt);
while (<$fh>) {
if (not $req_mark and /$trigger/) {
# Find the req_id of interest and save it into req_mark,
# then process the accumulated buffer
my ($req_id, $msg) = /:\s+(\*[0-9]+)\s+(.*)/;
$req_mark = $req_id;
# Find position of req_id which is skip_dist before the mark
# and print lines for req_mark before it (and all others)
my $del_idx = find_skip_start($req_mark, \@buf, $skip_dist);
for my $i (0..$#buf) {
if ($skip_idx and $i < $skip_idx) { print $buf[$i] }
else {
my ($req_id) = $buf[$i] =~ /:\s+(\*[0-9]+)/;
print $buf[$i] if $req_id ne $req_mark;
}
}
@buf = ();
$skip_idx = 0;
}
elsif (/$trigger/ or eof) {
# Process buffer collected between previous and this trigger,
# Or up to the end of file (the last line then need be added)
push @buf, $_ if eof;
my $skip_idx = (not eof)
? find_skip_start($req_mark, \@buf, $skip_dist)
: $#buf+1;
for my $i (0..$#buf) {
my ($req_id) = $buf[$i] =~ /:\s+(\*[0-9]+)/;
print $buf[$i]
if $req_id ne $req_mark
or (++$next_req_cnt > $skip_dist and $i < $skip_idx);
}
@buf = ();
$next_req_cnt = $skip_idx = 0;
# Check whether the request-id changed and update for next buffer
my ($req_id) = /:\s+(\*[0-9]+)/;
if ($req_id ne $req_mark) {
$req_mark = $req_id
}
}
else { push @buf, $_ }
}
sub find_skip_start {
my ($req_mark, $buf, $skip_dist) = @_;
my ($skip_idx, $prev_req_cnt);
for my $i (0..$#$buf) {
my ($req_id) = $buf->[$#$buf-$i] =~ /:\s+(\*[0-9]+)/;
if ( $req_id eq $req_mark and
(++$prev_req_cnt >= $skip_dist) )
{
$skip_idx = $#$buf-$i;
last;
}
}
return $skip_idx;
}
有许多地方可以提高效率。
一个主要的效率问题涉及缓冲区大小的中心问题。在我们扣动扳机之前能收集到多少数据?如果只有几个呢
[GET_..]
如果不知道这个短语的频率以及请求id的频率,就无法回答这个问题。然后一种优化方法是先浏览一下文件,估计一下频率,然后再做决定。但是,这并不太可靠,因为无法保证日志文件在任何意义上都是一致的。
上面的代码显然假设不会有太多的数据,因此我们不会造成麻烦,但是如果缓冲区太大,添加一个检查并写出缓冲区的一部分是明智的。
作为记录,当我对OP样本数据运行上面的程序时,输出是
2018/10/08 17:11:28 [debug] 8851#0: *2 Sent 8/8 bytes.
2018/10/08 17:11:28 [debug] 8851#0: *2 Session: Staging 8 bytes in thread buffer.
2018/10/08 17:11:33 [debug] 8851#0: *36 GET_REGION_INFO: Staging 99 bytes in thread buffer.
2018/10/08 17:11:33 [debug] 8851#0: *36 Sent 99/99 bytes.
2018/10/08 17:11:33 [debug] 8851#0: *36 Session: Staging 8 bytes in thread buffer.
2018/10/08 17:11:33 [debug] 8851#0: *36 Sent 8/8 bytes.
2018/10/08 17:11:38 [debug] 8851#0: *22 Receiving 8 bytes
2018/10/08 17:11:38 [debug] 8851#0: *22 Session: Staging 8 bytes in thread buffer.
[GET..]
(
$req_mark
)是
*36
.
*36
[获取..]
行)设置为两(2),以便更好地测试;这可以在调用时更改。
没有
*36
在输出的行之前加上
[GET...]
*36
之后
[获取..]
行(该行所在的位置)被省略,然后打印其余的行。这是预期的输出。
$skip_dist
)的
10
输出为
2018/10/08 17:11:28 [debug] 8851#0: *2 Sent 8/8 bytes.
2018/10/08 17:11:28 [debug] 8851#0: *2 Session: Staging 8 bytes in thread buffer.
2018/10/08 17:11:38 [debug] 8851#0: *22 Receiving 8 bytes
2018/10/08 17:11:38 [debug] 8851#0: *22 Session: Staging 8 bytes in thread buffer.
*36
在与
在这个样本数据中,所以没有
*36
原帖(相信
*36
是给定的感兴趣的请求id)
*36
缓存中的行,并在该区域中测试该短语。一旦你离开了那个区域,检查是否找到了这个短语并相应地打印出来
my $trigger = 'GET_REGION_INFO';
my $region_mark = '*36';
my (@buff, $drop_lines_mark);
while (<$fh>) {
my ($req_id, $msg) = /.*?:\s*(\*[0-9]+)\s+(.*)/;
if ($req_id eq $region_mark) {
push @buff, $_
$drop_lines_mark = $#buff if $msg =~ /$trigger/;
}
elsif (@buff) { # just left region of interest
if ($drop_lines_mark) {
for my $i (0..$#buff) {
print $buff[$i]
if $i < $drop_lines_mark-10
or $i > $drop_lines_mark+10;
}
}
else { print for @buff }
$drop_lines_mark = '';
@buff = ();
print; # don't forget the current line
}
else { print }
}
未测试的代码。