代码之家  ›  专栏  ›  技术社区  ›  Patrick Harrington

检测字符串中的URL,并用“<a href…”标签包裹

  •  10
  • Patrick Harrington  · 技术社区  · 17 年前

    我想写一些看起来很容易的东西,但无论出于什么原因,我都很难理解。

    unencoded_string = "This is a link - http://google.com"
    
    def encode_string_with_links(unencoded_string):
        # some sort of regex magic occurs
        return encoded_string
    
    print encoded_string
    
    'This is a link - <a href="http://google.com">http://google.com</a>'
    

    非常感谢。

    2 回复  |  直到 12 年前
        1
  •  11
  •   Laurence Gonsalves    17 年前

    你需要的“正则表达式魔法”只是 sub

    def encode_string_with_links(unencoded_string):
      return URL_REGEX.sub(r'<a href="\1">\1</a>', unencoded_string)
    

    URL_REGEX 可能类似于:

    URL_REGEX = re.compile(r'''((?:mailto:|ftp://|http://)[^ <>'"{}|\\^`[\]]*)''')
    

        2
  •  11
  •   Keyrr Perino tefozi    8 年前

    谷歌搜索解决方案:

    #---------- find_urls.py----------#
    # Functions to identify and extract URLs and email addresses
    
    import re
    
    def fix_urls(text):
        pat_url = re.compile(  r'''
                         (?x)( # verbose identify URLs within text
             (http|ftp|gopher) # make sure we find a resource type
                           :// # ...needs to be followed by colon-slash-slash
                (\w+[:.]?){2,} # at least two domain groups, e.g. (gnosis.)(cx)
                          (/?| # could be just the domain name (maybe w/ slash)
                    [^ \n\r"]+ # or stuff then space, newline, tab, quote
                        [\w/]) # resource name ends in alphanumeric or slash
             (?=[\s\.,>)'"\]]) # assert: followed by white or clause ending
                             ) # end of match group
                               ''')
        pat_email = re.compile(r'''
                        (?xm)  # verbose identify URLs in text (and multiline)
                     (?=^.{11} # Mail header matcher
             (?<!Message-ID:|  # rule out Message-ID's as best possible
                 In-Reply-To)) # ...and also In-Reply-To
                        (.*?)( # must grab to email to allow prior lookbehind
            ([A-Za-z0-9-]+\.)? # maybe an initial part: DAVID.mertz@gnosis.cx
                 [A-Za-z0-9-]+ # definitely some local user: MERTZ@gnosis.cx
                             @ # ...needs an at sign in the middle
                  (\w+\.?){2,} # at least two domain groups, e.g. (gnosis.)(cx)
             (?=[\s\.,>)'"\]]) # assert: followed by white or clause ending
                             ) # end of match group
                               ''')
    
        for url in re.findall(pat_url, text):
           text = text.replace(url[0], '<a href="%(url)s">%(url)s</a>' % {"url" : url[0]})
    
        for email in re.findall(pat_email, text):
           text = text.replace(email[1], '<a href="mailto:%(email)s">%(email)s</a>' % {"email" : email[1]})
    
        return text
    
    if __name__ == '__main__':
        print fix_urls("test http://google.com asdasdasd some more text")
    

    编辑: 根据您的需求进行调整