代码之家  ›  专栏  ›  技术社区  ›  Daniel Haley

如何使用lxml选择和更新混合内容中的文本节点?

  •  0
  • Daniel Haley  · 技术社区  · 7 年前

    我需要检查所有的单词 text() XML文件中的节点。我正在使用XPath //text() 选择文本节点,使用正则表达式选择单词。如果这个词存在于一组关键字中,我需要用一些东西替换它并更新XML。

    通常,设置元素的文本是使用 .text 但是 文本 在_元素上,只会更改第一个子文本节点。在一个 mixed content element ,其他文本节点实际上是 .tail 它的前兄弟姐妹。

    如何更新所有文本节点?

    在下面的简化示例中,我只是尝试将匹配的关键字用方括号括起来。。。

    输入XML

    <doc>
        <para>I think the only card she has <gotcha>is the</gotcha> Lorem card. We have so many things that we have to do
            better... and certainly ipsum is one of them. When other <gotcha>websites</gotcha> give you text, they're not
            sending the best. They're not sending you, they're <gotcha>sending words</gotcha> that have lots of problems
            and they're <gotcha>bringing</gotcha> those problems with us. They're bringing mistakes. They're bringing
            misspellings. They're typists… And some, <gotcha>I assume</gotcha>, are good words.</para>
    </doc>
    

    期望输出

    <doc>
        <para>I think [the] only card she has <gotcha>[is] [the]</gotcha> Lorem card. We have so many things that we have to do
            better... and certainly [ipsum] [is] one of them. When other <gotcha>websites</gotcha> give you text, they're not
            sending [the] [best]. They're not sending you, they're <gotcha>sending words</gotcha> that have lots of [problems]
            and they're <gotcha>bringing</gotcha> those [problems] with us. They're bringing [mistakes]. They're bringing
            misspellings. They're typists… And some, <gotcha>I assume</gotcha>, are good words.</para>
    </doc>
    
    1 回复  |  直到 7 年前
        1
  •  3
  •   Daniel Haley    7 年前

    我在文档中找到了这个解决方案的关键: Using XPath to find text

    特别是 is_text 和 is_tail 特性 _ElementUnicodeResult .

    使用这些属性,我可以判断是否需要更新 .text 或 .tail 父母的财产 _Element .

    这是一个有点棘手的理解,因为当你使用 getparent() 在文本节点上( _ElementUnicodeResult )这是它之前兄弟的尾巴( .is_tail == True ),前一个兄弟姐妹作为父代返回;不是真正的父母。

    实例

    python

    import re
    from lxml import etree
    
    xml = """<doc>
        <para>I think the only card she has <gotcha>is the</gotcha> Lorem card. We have so many things that we have to do
            better... and certainly ipsum is one of them. When other <gotcha>websites</gotcha> give you text, they're not
            sending the best. They're not sending you, they're <gotcha>sending words</gotcha> that have lots of problems
            and they're <gotcha>bringing</gotcha> those problems with us. They're bringing mistakes. They're bringing
            misspellings. They're typists… And some, <gotcha>I assume</gotcha>, are good words.</para>
    </doc>
    """
    
    
    def update_text(match, word_list):
        if match in word_list:
            return f"[{match}]"
        else:
            return match
    
    
    root = etree.fromstring(xml)
    
    keywords = {"ipsum", "is", "the", "best", "problems", "mistakes"}
    
    for text in root.xpath("//text()"):
        parent = text.getparent()
        updated_text = re.sub(r"[\w]+", lambda match: update_text(match.group(), keywords), text)
        if text.is_text:
            parent.text = updated_text
        elif text.is_tail:
            parent.tail = updated_text
    
    etree.dump(root)
    

    输出 (转储到控制台)

    <doc>
        <para>I think [the] only card she has <gotcha>[is] [the]</gotcha> Lorem card. We have so many things that we have to do
            better... and certainly [ipsum] [is] one of them. When other <gotcha>websites</gotcha> give you text, they're not
            sending [the] [best]. They're not sending you, they're <gotcha>sending words</gotcha> that have lots of [problems]
            and they're <gotcha>bringing</gotcha> those [problems] with us. They're bringing [mistakes]. They're bringing
            misspellings. They're typists… And some, <gotcha>I assume</gotcha>, are good words.</para>
    </doc>