代码之家  ›  专栏  ›  技术社区  ›  Tiger1

如何在不指定群集数量的情况下对列表中的项目进行群集

  •  2
  • Tiger1  · 技术社区  · 12 年前

    我的目标是根据任意两个连续项目之间的距离对以下列表中的项目进行聚类。未指定簇数,仅指定最大距离 在任何两个连续项目之间,如果要使它们位于同一集群中,则不能超过该值。

    我的尝试

    import itertools
    max_dist=20
    _list=[1,48,52,59,89,94,103,147,151,165]
    Ideal_result= [[1],[48,52,59],[89,94,103],[147,151,165]]
    
    def clust(list_x, max_dist):
        q=[]
        for a, b in itertools.combinations(list_x,2):
            if b-a<=20:
                q.append(a),(b)
            else:continue
            yield q
    print list(clust(_list,max_dist))
    

    输出:

    [[48,48,52,89,89,94,147,147,151],[48,48,52,89,89,94,147,147,151],..]`
    

    输出完全错误,但我只是想包括我的尝试。

    对如何获得理想结果有什么建议吗?谢谢

    1 回复  |  直到 12 年前
        1
  •  2
  •   jscs    12 年前

    这通过了您的测试:

    def cluster(items, key_func):
        items = sorted(items)
        clusters = [[items[0]]]
        for item in items[1:]:
            cluster = clusters[-1]
            last_item = cluster[-1]
            if key_func(item, last_item):
                cluster.append(item)
            else:
                clusters.append([item])
        return clusters
    

    哪里 key_func 回报 True 如果当前项目和先前项目应属于同一集群:

    >>> cluster([1,48,52,59,89,94,103,147,151,165], lambda curr, prev: curr-prev < 20)
    [[1], [48, 52, 59], [89, 94, 103], [147, 151, 165]]
    

    另一种可能是修改 "equivalent code" for itertools.groupby() 同样为键函数取多个参数。