代码之家  ›  专栏  ›  技术社区  ›  Iain Samuel McLean Elder

HtmlUnit和XPath:DOMNode.getByXPath只适用于HtmlPage?

  •  1
  • Iain Samuel McLean Elder  · 技术社区  · 15 年前

    我在分析 a page 链接到重要内容如下所示的文章:

    <div class="article">
      <h1 style="float: none;"><a href="performing-arts">Performing Arts</a></h1>
      <a href="/performing-arts/EIF-theatre-review-Sin-Sangre.6517348.jp">
        <span class="mth3">
          <span id="wctlMiniTemplate1_ctl00_ctl00_ctl01_WctlPremiumContentIcon1">                               
          </span>
          EIF theatre review: Sin Sangre | The Man Who Fed Butterflies | Caledonia | Songs Of Ascension | Vieux Carré | The Gospel At Colonus
        </span>
        <span class="mtp">The EIF&#39;s theatre programme wasn&#39;t as far-reaching as it could have been, but did find an exoticism in the familiar,  writes Mark Fisher </span>
      </a>                  
    </div>
    

    下面是Java中使用HtmlUnit和XPath的一个最小刮取案例(为简洁起见,删除了导入):

    public class MinimalTest {
        public static void main(String[] args) throws Exception {
            WebClient client = new WebClient();
            client.setJavaScriptEnabled(false);
            client.setCssEnabled(false);
            System.out.println("Fetching front page");
            HtmlPage frontPage = client.getPage("http://living.scotsman.com/sectionhome.aspx?sectionID=7063");
            List<ArticleInfo> articleInfos = extractArticleInfo(frontPage);
    
            for (ArticleInfo info : articleInfos)
            {
                System.out.println("Title: " + info.getTitle());
                System.out.println("Intro: " + info.getFirstPara());
                System.out.println("Link: " + info.getLink());
            }
        }
    
        @SuppressWarnings("unchecked") // xpath returns List<?>
        private static List<ArticleInfo> extractArticleInfo(HtmlPage frontPage) {
            System.out.println("Extracting article links");
            List<HtmlDivision> articleDivs = (List<HtmlDivision>) frontPage.getByXPath("//div[@class='article']");
            System.out.println(String.format("Found %d articles", articleDivs.size()));
            List<ArticleInfo> articleLinks = new ArrayList<ArticleInfo>(articleDivs.size());
            for (HtmlDivision div : articleDivs) {
                articleLinks.add(ArticleInfo.constructFromArticleDiv(div));
            }
            return articleLinks;
        }
    
        private static class ArticleInfo {
            private final String title;
            private final String link;
            private final String firstPara;
    
            public ArticleInfo(final String link, final String title, final String firstPara) {
                this.link = link;
                this.title = title;
                this.firstPara = firstPara;
            }
            public static ArticleInfo constructFromArticleDiv(final HtmlDivision div) {
                String link = ((DomText) div.getFirstByXPath("//a/@href/text()")).asText();
                String title = ((DomText) div.getFirstByXPath("//span[@class='mth3']/text()")).asText();
                String firstPara = ((DomText) div.getFirstByXPath("//span[@class='mtp']/text()")).asText();
                return new ArticleInfo(link, title, firstPara);
            }
            public String getTitle() {
                return title;
            }
            public String getFirstPara() {
                return firstPara;
            }
            public String getLink() {
                return link;
            }
        }  
    }
    

    我期望的输出:

    Title: EIF theatre review: Sin Sangre | The Man Who Fed Butterflies | Caledonia | Songs Of Ascension | Vieux Carré | The Gospel At Colonus 
    Intro: The EIF's theatre programme wasn't as far-reaching as it could have been, but did find an exoticism in the familiar, writes Mark Fisher 
    Link: http://living.scotsman.com/performing-arts/EIF-theatre-review-Sin-Sangre.6517348.jp
    

    我得到的是:

    Fetching front page
    Extracting article links
    Found 24 articles
    Exception in thread "main" java.lang.NullPointerException
        at com.allthefestivals.app.crawler.MinimalTest$ArticleInfo.constructFromArticleDiv(MinimalTest.java:68)
        at com.allthefestivals.app.crawler.MinimalTest.extractArticleInfo(MinimalTest.java:50)
        at com.allthefestivals.app.crawler.MinimalTest.main(MinimalTest.java:30)
        at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)
        at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:39)
        at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25)
        at java.lang.reflect.Method.invoke(Method.java:597)
        at com.intellij.rt.execution.application.AppMain.main(AppMain.java:115)
    

    getByXPath 在一个地方很好用 HtmlPage 但似乎没有任何回报 HtmlElement

    相关问题的解决方案对我不起作用: XPath _relative_ to given element in HTMLUnit/Groovy?

    1 回复  |  直到 9 年前
        1
  •  2
  •   Rodney Gitzel    15 年前

    您已尝试将属性视为元素。请尝试以下操作:

    String link = ((DomAttr) div.getFirstByXPath("//a/@href")).getValue();
    

    Fetching front page
    Extracting article links
    Found 24 articles
    Title: EIF theatre review: Sin Sangre | The Man Who Fed Butterflies | Caledonia | Songs Of Ascension | Vieux Carré | The Gospel At Colonus
    Intro: The EIF's theatre programme wasn't as far-reaching as it could have been, but did find an exoticism in the familiar, writes Mark Fisher
    Link: /Register.aspx?ReturnURL=http%3a%2f%2fliving.scotsman.com%2fsectionhome.aspx%3fsectionID%3d7063
    ...
    

    另外,ArticleInfo类将“link”声明为字符串,然后为其赋值(custom?)班级。为了让它编译,我不得不把东西弄乱一点。

    推荐文章