我的HTML如下所示。
<div class="topics">
<h2>Topics 1</h2>
News
Sports
<h2>Topics 2</h2>
Entertainment
Business
</div>
我希望能够得到文本
["News\nSports", "Entertainment\nBusiness"]
使用
XPath
.我该怎么做?
//div[contains(@class,"topics")]/h2/text()
给了我
h2
文本,但我也想要下面相应的(以下)文本。
//div[contains(@class,"topics")]/h2/following-sibling::text()
之后给我所有的文字吗
但是以这样的方式排列
["News", "\n", "Sports", "Entertainment", "\n", "Business"]
.现在我无法将文本字符串数组与标题相关联。
我在用
Scrapy v1.5.1
发出XPath。
content.xpath("//div[contains(@class,"topics")]/h2/following-sibling::text()").extract()
奇怪的是,这个XPath查询在Chrome中工作(查看黄色高亮显示的文本),但不是通过Scrapy。