代码之家  ›  专栏  ›  技术社区  ›  Vincent

使用lxml解析文本,并使用一些标记将其分解为句子列表以添加结构

  •  0
  • Vincent  · 技术社区  · 3 年前

    在自定义中考虑以下文本 xml :

    <?xml version="1.0"?>
    <body>
        <heading><b>This is a title</b></heading>
        <p>This is a first <b>paragraph</b>.</p>
        <p>This is a second <b>paragraph</b>. With a list: 
            <ul>
                <li>first item</li>
                <li>second item</li>
            </ul>
        And the end.
        </p>
        <p>This is a third paragraph.
            <ul>
                <li>This is a first long sentence.</li>
                <li>This is a second long sentence.</li>
            </ul>
        And the end of the paragraph.</p>
    </body>
    

    我想用以下规则将其转换为纯字符串列表:

    • 放弃一些标签,如 <b></b>
    • 每个 heading 和每个 paragraph 是列表中不同的元素。如果元素末尾缺少最后一个句点,则添加该句点。
    • 当列表前面有冒号“:”时,只需在元素之间添加换行符并添加短划线即可。
    • 如果列表前面没有冒号,则将该段落拆分为多个段落

    结果是:

    [
        "This is a title.", # Note the period
        "This is a first paragraph.",
        "This is a second paragraph. With a list:\n- first item\n- second item\nAnd the end.",
        "This is a third paragraph.",
        "This is a first long sentence.",
        "This is a second long sentence.",
        "And the end of the paragraph."
    ]
    

    我想通过迭代lmxl etree的结果来实现这一点 etree.fromstring(text) 。我的前几次试验过于复杂和缓慢,我相信有一个很好的方法来解决这个问题。

    怎么做?

    0 回复  |  直到 3 年前
        1
  •  2
  •   Jack Fleeting    3 年前

    有趣的运动。。。

    以下内容有点复杂,不会给你所指示的确切输出,但可能它足够接近,你(或其他人)可以修改它:

    from lxml import etree
    stuff = """[your xml]"""
            
    doc =  etree.XML(stuff)
        
    #we need this in order to count how many <li> elements meet the condition
    #in your xml there are only two, but this will take care of more elements
    comms = len(doc.xpath('//p[contains(.,":")]//ul//li'))
    final = []
        
    for t in doc.xpath('//*'):
        line = "".join(list(t.itertext()))    
        allin = [l.strip() for l in line.split('\n  ') if len(l.strip())>0]
        for l in allin:
            ind = allin.index(l)
            for c in range(comms):
                if ":" in allin[ind-(c+1)]:
                    final.append("- "+l)
            if l[-1] =="." or l[-1] ==":":
                final.append(l)
            else:
                if not ("- "+l in final):
                    final.append(l+".")
        break
     
    final
    

    输出:

    ['This is a title.',
     'This is a first paragraph.',
     'This is a second paragraph. With a list:',
     '- first item',
     '- second item',
     'And the end.',
     'This is a third paragraph.',
     'This is a first long sentence.',
     'This is a second long sentence.',
     'And the end of the paragraph.']