代码之家  ›  专栏  ›  技术社区  ›  Ray

在PHP代码库中查找所有字符串

  •  4
  • Ray  · 技术社区  · 17 年前

    我有几百万行PHP代码库 没有 真正的显示和逻辑分离,为了本地化,我试图提取代码中表示的所有字符串。将显示和逻辑分离是一个长期目标,但目前我只希望能够本地化。

    在代码中,字符串以PHP的各种可能格式表示,因此我需要一种理论(或实际)方法来解析整个源代码,至少找到每个字符串所在的位置。当然,在理想情况下,我会用一个函数调用替换每个字符串

    "this is a string"

    将被替换为

    _("this is a string")

    当然,我需要同时支持单人和双人 quote format .其他我不太关心的,它们很少出现,我可以手动更改它们。

    当然,我也不想本地化数组索引。所以弦像

    $arr["value"]

    不应该成为

    $arr[_("value")]

    有人能帮我开始吗?

    3 回复  |  直到 17 年前
        1
  •  11
  •   Tom Haigh    17 年前

    你可以用 token_get_all() 从PHP文件中获取所有令牌 例如

    <?php
    
    $fileStr = file_get_contents('file.php');
    
    foreach (token_get_all($fileStr) as $token) {
        if ($token[0] == T_CONSTANT_ENCAPSED_STRING) {
            echo "found string {$token[1]}\r\n";
            //$token[2] is line number of the string
        }
    }
    

    你可以做一个非常糟糕的检查,确保它没有被用作数组索引,比如:

    $fileLines = file('file.php');
    
    //inside the loop and if
    $line = $fileLines[$token[2] - 1];
    if (false === strpos($line, "[{$token[1]}]")) {
        //not an array index
    }
    

    但你真的很难做到这一点,因为可能有人写了一些你可能并不期待的东西,例如:

    $str = 'string that is not immediately an array index';
    doSomething($array[$str]);
    

    编辑 正如Ant P所说,你可能会更好地寻找 [ 和 ] 在这个答案的第二部分,而不是我的周围标记 strpos 哈克,类似这样的:

    $i = 0;
    $tokens = token_get_all(file_get_contents('file.php'));
    $num = count($tokens);
    for ($i = 0; $i < $num; $i++) {
        $token = $tokens[$i];
    
        if ($token[0] != T_CONSTANT_ENCAPSED_STRING) {
            //not a string, ignore
            continue;
        }
    
        if ($tokens[$i - 1] == '[' && $tokens[$i + 1] == ']') {
            //immediately used as an array index, ignore
            continue; 
        }
    
        echo "found string {$token[1]}\r\n";
        //$token[2] is line number of the string
    }
    
        2
  •  5
  •   postfuturist    17 年前

    在代码库中还可能存在一些其他情况,除了关联数组之外,您还可以通过执行自动搜索和替换来彻底打破这些情况。

    SQL查询:

    $myname = "steve";
    $sql = "SELECT foo FROM bar WHERE name = " . $myname;
    

    间接变量引用。

    $bar = "Hello, World"; // a string that needs localization
    $foo = "bar"; // a string that should not be localized
    echo($$foo);
    

    SQL字符串操作。

    $sql = "SELECT CONCAT('Greetings, ', firstname) as greeting from users where id = ?";
    

    没有自动过滤所有可能性的方法。也许解决方案是编写一个应用程序,创建一个可能字符串的“调节”队列,并在几行代码的上下文中显示每个突出显示的字符串。然后,您可以浏览代码以确定它是否是一个需要本地化的字符串,然后点击一个键来本地化或忽略该字符串。

        3
  •  -3
  •   brettkelly    17 年前

    不要试图通过使用perl或grep进行过于聪明的命令行攻击来解决这个问题,而应该编写一个程序来实现这一点:)

    编写一个perl/python/ruby/anywhere脚本,在每个文件中搜索一对单引号或双引号。每次它找到匹配项时,都会提示您用下划线函数替换它,您可以告诉它这样做,也可以跳到下一个。

    在一个完美的世界里,你会写一些能帮你完成这一切的东西,但这最终可能会花更少的时间,你会面临更少的错误。

    伪:

    for fname in yourBigFileList:
        create file handle for actual source file
        create temp file handle (like fname +".tmp" or something)
        for fline in fname:
            get quoted strings
            for qstring in quoted_strings:
                show it in context, i.e. the entire line of code.
                replace with _()?
                    if Y, replace and write line to tmp file
                    if N, just write that line to the tmp file
        close file handles
        rename it to current name + ".old"
        rename ".tmp" file to name of orignal file
    

    我相信有一种更简单的方法可以做到这一点,但这种方法可以让你自己查看每个实例并做出决定。如果有一百万行,每一行都包含一个字符串,每一行需要你1秒的时间来计算,那么你需要270个小时来完成整个过程。。。也许你应该忽略这个帖子:)