代码之家  ›  专栏  ›  技术社区  ›  fifigyuri

如何使用wildchards,模糊搜索与Solr?

  •  3
  • fifigyuri  · 技术社区  · 16 年前

    我使用Solr搜索我的数据,现在我意识到一些Solr搜索查询语言特性并不适合我。我很怀念我的能力:

    • 模糊搜索
    • 字段规格-目前我无法告诉搜索标题:blablabla

    据我所知,这些东西在Solr中应该是默认的,但是我显然没有它们。我用solr1.4。在这里你可以找到 my schema . 谢谢你的帮助。

    2 回复  |  直到 16 年前
        1
  •  5
  •   Brain    13 年前

    我在谷歌上搜索“ “我在这里找到了你的问题。实际上,SOLR的4.0版本能够使用简单的查询语法进行模糊搜索。

    例如,您可以搜索 name:peter 严格的或带有颚化符号的 name:peter~ 作为模糊搜索。如果你能用一点模糊的形式来限制欲望的话 name:peter~0.7 锐度 “70%。

        2
  •  4
  •   Mauricio Scheffer    16 年前

    你的 fieldType name="text"

    <!-- A text field that uses WordDelimiterFilter to enable splitting and matching of
        words on case-change, alpha numeric boundaries, and non-alphanumeric chars,
        so that a query of "wifi" or "wi fi" could match a document containing "Wi-Fi".
        Synonyms and stopwords are customized by external files, and stemming is enabled.
        -->
    <fieldType name="text" class="solr.TextField" positionIncrementGap="100">
      <analyzer type="index">
        <tokenizer class="solr.WhitespaceTokenizerFactory"/>
        <!-- in this example, we will only use synonyms at query time
        <filter class="solr.SynonymFilterFactory" synonyms="index_synonyms.txt" ignoreCase="true" expand="false"/>
        -->
        <!-- Case insensitive stop word removal.
          add enablePositionIncrements=true in both the index and query
          analyzers to leave a 'gap' for more accurate phrase queries.
        -->
        <filter class="solr.StopFilterFactory"
                ignoreCase="true"
                words="stopwords.txt"
                enablePositionIncrements="true"
                />
        <filter class="solr.WordDelimiterFilterFactory" generateWordParts="1" generateNumberParts="1" catenateWords="1" catenateNumbers="1" catenateAll="0" splitOnCaseChange="1"/>
        <filter class="solr.LowerCaseFilterFactory"/>
        <filter class="solr.SnowballPorterFilterFactory" language="English" protected="protwords.txt"/>
      </analyzer>
      <analyzer type="query">
        <tokenizer class="solr.WhitespaceTokenizerFactory"/>
        <filter class="solr.SynonymFilterFactory" synonyms="synonyms.txt" ignoreCase="true" expand="true"/>
        <filter class="solr.StopFilterFactory"
                ignoreCase="true"
                words="stopwords.txt"
                enablePositionIncrements="true"
                />
        <filter class="solr.WordDelimiterFilterFactory" generateWordParts="1" generateNumberParts="1" catenateWords="0" catenateNumbers="0" catenateAll="0" splitOnCaseChange="1"/>
        <filter class="solr.LowerCaseFilterFactory"/>
        <filter class="solr.SnowballPorterFilterFactory" language="English" protected="protwords.txt"/>
      </analyzer>
    </fieldType>
    

    例如 SnowballPorterFilterFactory 是启用词干分析的。

    我建议根据默认模式构建架构.xml,根据需要进行调整和修改(而不是从头开始)。

    Here's the reference for analyzers, tokenizers and filters .