代码之家  ›  专栏  ›  技术社区  ›  user3525290

使用BeautifulSoup将HTML插入元素

  •  1
  • user3525290  · 技术社区  · 8 年前

    当我尝试将以下HTML插入到元素中时

    <div class="frontpageclass"><h3 id="feature_title">The Title</h3>... </div>
    

    bs4 是这样替换的:

    <div class="frontpageclass">&lt;h3 id="feature_title"&gt;The Title &lt;/h3&gt;... &lt;div&gt;</div>
    

    string 它仍然在搞砸格式。

    with open(html_frontpage) as fp:
       soup = BeautifulSoup(fp,"html.parser")
    
    found_data = soup.find(class_= 'front-page__feature-image')
    found_data.string = databasedata
    

    如果我尝试使用 found_data.string.replace_with 我得到一个非类型错误。 found_data 类型为tag。

    similar issue but they are using div, not class

    1 回复  |  直到 8 年前
        1
  •  5
  •   Tomalak    8 年前

    设置元素 .text .string

    实际的 HTML,您需要在树中插入新节点。

    from bs4 import BeautifulSoup
    
    # always define a file encoding when working with text files
    with open(html_frontpage, encoding='utf8') as fp:
        soup = BeautifulSoup(fp, "html.parser")
    
    target = soup.find(class_= 'front-page__feature-image')
    
    # empty out the target element if needed
    target.clear()
    
    # create a temporary document from your HTML
    content = '<div class="frontpageclass"><h3 id="feature_title">The Title</h3>...</div>'
    temp = BeautifulSoup(content)
    
    # the nodes we want to insert are children of the <body> in `temp`
    nodes_to_insert = temp.find('body').children
    
    # insert them, in source order
    for i, node in enumerate(nodes_to_insert):
        target.insert(i, node)
    
        2
  •  0
  •   Fanghaozhi Daniel Liang    6 年前

    对于messing格式,只有与“<”和“>”对应的<和>。只要把它们全部换掉就行了。

    例如,假设beautifulsoup以混乱的格式将html标记插入soup1变量: a=str(soup1).replace(&lt;,'<').replace(&gt;,'>');print(a)

    在实际代码中,应该将<放在“”内,中间不留空格。(此处,web显示<与<没有相同的空间)

    所以变量a应该使用正确的格式。