代码之家  ›  专栏  ›  技术社区  ›  torger

如何使用lxml、xpath和python从网页中提取链接?

  •  5
  • torger  · 技术社区  · 16 年前

    我有一个xpath查询:

    /html/body//tbody/tr[*]/td[*]/a[@title]/@href
    

    它提取具有标题属性的所有链接,并给出 href 在里面 FireFox's Xpath checker add-on .

    但是,我似乎不能用它 lxml .

    from lxml import etree
    parsedPage = etree.HTML(page) # Create parse tree from valid page.
    
    # Xpath query
    hyperlinks = parsedPage.xpath("/html/body//tbody/tr[*]/td[*]/a[@title]/@href") 
    for x in hyperlinks:
        print x # Print links in <a> tags, containing the title attribute
    

    这不会导致 LXML (空列表)。

    怎样才能抓住 HREF 包含属性标题的超链接的文本(链接) LXML 在蟒蛇下面?

    2 回复  |  直到 13 年前
        1
  •  10
  •   jkp    16 年前

    我可以使用以下代码:

    from lxml import html, etree
    from StringIO import StringIO
    
    html_string = '''<!DOCTYPE html PUBLIC "-//W3C//DTD HTML 4.01 Transitional//EN"
       "http://www.w3.org/TR/html4/loose.dtd">
    
    <html lang="en">
    <head/>
    <body>
        <table border="1">
          <tbody>
            <tr>
              <td><a href="http://stackoverflow.com/foobar" title="Foobar">A link</a></td>
            </tr>
            <tr>
              <td><a href="http://stackoverflow.com/baz" title="Baz">Another link</a></td>
            </tr>
          </tbody>
        </table>
    </body>
    </html>'''
    
    tree = etree.parse(StringIO(html_string))
    print tree.xpath('/html/body//tbody/tr/td/a[@title]/@href')
    
    >>> ['http://stackoverflow.com/foobar', 'http://stackoverflow.com/baz']
    
        2
  •  2
  •   mrmagooey    14 年前

    火狐 adds additional html tags 在呈现HTML时,使Firebug工具返回的xpath与服务器返回的实际HTML不一致(以及urllib/2将返回的内容)。

    去掉 <tbody> 标签通常起作用。

    推荐文章