代码之家  ›  专栏  ›  技术社区  ›  Costantin

使用regex和javascript从字符串中获取基URL

  •  0
  • Costantin  · 技术社区  · 7 年前

    我试图从一个字符串中获取基URL(所以没有window.location)。

    • 它需要删除尾随斜线
    • 它需要是regex(没有新的url)
    • 它需要使用查询参数和锚链接

    换言之,以下所有内容都应返回 https://apple.com 或 https://www.apple.com 最后一个。

    • https://apple.com?query=true&slash=false
    • https://apple.com#anchor=true&slash=false
    • http://www.apple.com/#anchor=true&slash=true&whatever=foo

    这些只是示例,URL可以具有不同的子域,例如 https://shop.apple.co.uk/?query=foo 应该返回 https://shop.apple.co.uk -它可以是任何URL,比如: https://foo.bar

    我越接近:

    const baseUrl = url.replace(/^((\w+:)?\/\/[^\/]+\/?).*$/,'$1').replace(/\/$/, ""); // Base Path & Trailing slash
    

    但这不适用于定位链接和查询,这些链接和查询在没有 / 之前

    我知道如何让它处理所有案件吗?

    5 回复  |  直到 7 年前
        1
  •  1
  •   The fourth bird    7 年前

    你可以添加 # 和 ? 对你 negated character class . 你不需要 .* 因为这将匹配到字符串的末尾。

    对于示例数据,可以 match :

    ^https?:\/\/[^#?\/]+
    

    Regex demo

    strings = [
    "https://apple.com?query=true&slash=false",
        "https://apple.com#anchor=true&slash=false",
        "http://www.apple.com/#anchor=true&slash=true&whatever=foo",
        "https://foo.bar/?q=true"
    ];
    
    strings.forEach(s => {
        console.log(s.match(/^https?:\/\/[^#?\/]+/)[0]);
    })
        2
  •  0
  •   sui    7 年前
        const baseUrl = url.replace(/(.*:\/\/.*)[\?\/#].*/, '$1');
    
        3
  •  0
  •   Matt Morgan    7 年前

    你 能够 使用javascript的内置 URL 为了这个。URL还将为您提供其他易于访问的已解析属性,如查询字符串参数、协议等。

    regex是一种痛苦的方式,它可以让javascript做一些其他方面非常简单的事情。

    我知道你问过关于使用regex的问题,但是如果你(或将来来这里的人)真的只关心获取信息,而不致力于使用regex,也许这个答案会有所帮助。

    let one = "https://apple.com?query=true&slash=false"
    let two = "https://apple.com#anchor=true&slash=false"
    let three = "http://www.apple.com/#anchor=true&slash=true&whatever=foo"
    
    let urlOne = new URL(one)
    console.log(urlOne.origin)
    
    let urlTwo = new URL(two)
    console.log(urlTwo.origin)
    
    let urlThree = new URL(three)
    console.log(urlThree.origin)
        4
  •  0
  •   user875234    7 年前

    这将使您的所有内容都在.com部分。一旦拉出URL的第一部分,就必须附加.com。

    ^http.*?(?=\.com)
    

    或者你可以这样做:

    myUrl.Replace(/(#|\?|\/#).*$/, "")
    

    删除主机名后的所有内容。

        5
  •  -1
  •   samnu pel    7 年前

    你可以这样做

    if(url.indexOf('#') !== -1) { var baseUrl = url.split("#")[0]; } else if (url.indexOf('?') !== -1) { var baseUrl = url.split("?")[0]; } else { var baseUrl =  url }