代码之家  ›  专栏  ›  技术社区  ›  hbstha123

如何从Python中的pandas数据帧中获得networkx图的分支作为列表?

  •  0
  • hbstha123  · 技术社区  · 4 年前

    我有一个熊猫数据帧 df 其看起来如下:

    From    To
    0   Node1   Node2
    1   Node1   Node3
    2   Node2   Node4
    3   Node2   Node5
    4   Node3   Node6
    5   Node3   Node7
    6   Node4   Node8
    7   Node5   Node9
    8   Node6   Node10
    9   Node7   Node11
    

    df.to_dict() 是

    {'From': {0: 'Node1',
      1: 'Node1',
      2: 'Node2',
      3: 'Node2',
      4: 'Node3',
      5: 'Node3',
      6: 'Node4',
      7: 'Node5',
      8: 'Node6',
      9: 'Node7'},
     'To': {0: 'Node2',
      1: 'Node3',
      2: 'Node4',
      3: 'Node5',
      4: 'Node6',
      5: 'Node7',
      6: 'Node8',
      7: 'Node9',
      8: 'Node10',
      9: 'Node11'}}
    

    我使用networkx软件包将熊猫数据帧绘制为网络图,如下所示: enter image description here

    我想从这个网络图中获得唯一场景/分支的列表。 从Node1开始,这里有四个分支。

    Node1-Node2-Node4-Node8
    Node1-Node2-Node5-Node9
    Node1-Node3-Node6-Node10
    Node1-Node3-Node7-Node11
    

    如何从Python中给定的pandas数据帧中获得上面的分支列表?

    0 回复  |  直到 4 年前
        1
  •  2
  •   I'mahdi    4 年前

    您可以定义 Recursive Function 以及保存路径和打印路径:

    df = pd.DataFrame({
              'From':['Node1','Node1', 'Node2', 'Node2', 'Node3', 'Node3', 'Node4', 'Node5', 'Node6', 'Node7'],
              'TO'  :['Node2','Node3', 'Node4', 'Node5', 'Node6', 'Node7', 'Node8', 'Node9', 'Node10', 'Node11']
            })
    
    def prntPath(lst, node, df, lst_vst):
        for val in df.values:
            if val[0] == node:
                lst.append(val[1])
                prntPath(lst, val[1], df, lst_vst)
        
        if not lst[-1] in lst_vst:
            print('-'.join(lst))
        for l in lst: lst_vst.add(l)
        lst.pop()
        return
        
    lst_vst = set()
    prntPath(['Node1'],'Node1', df, lst_vst)
    

    输出

    Node1-Node2-Node4-Node8
    Node1-Node2-Node5-Node9
    Node1-Node3-Node6-Node10
    Node1-Node3-Node7-Node11
    

    您可以检查并用于其他图形,如下所示:

    import networkx as nx
    import matplotlib.pyplot as plt
    import pandas as pd
    from itertools import chain
    from networkx.drawing.nx_pydot import graphviz_layout
    
    def prntPath(lst, node, df, lst_vst):
        for val in df.values:
            if val[0] == node:
                lst.append(val[1])
                prntPath(lst, val[1], df, lst_vst)
        if not lst[-1] in lst_vst: print('-'.join(lst))
        for l in lst: lst_vst.add(l)
        lst.pop()
        return
    
    df = pd.DataFrame({
              'From':['Node1','Node1', 'Node2', 'Node3', 'Node3', 'Node5', 'Node7'],
              'TO'  :['Node2','Node3', 'Node5', 'Node6', 'Node7', 'Node9', 'Node11']
            })
    
    g = nx.DiGraph()
    g.add_nodes_from(set(chain.from_iterable(df.values)))
    for edg in df.values:
        g.add_edge(*edg)
    pos = graphviz_layout(g, prog="dot")
    nx.draw(g, pos,with_labels=True, node_shape='s')
    plt.draw()
    plt.show() 
    
    lst_vst = set()
    prntPath(['Node1'],'Node1', df, lst_vst)
    

    输出

    enter image description here

    Node1-Node2-Node5-Node9
    Node1-Node3-Node6
    Node1-Node3-Node7-Node11
    
        2
  •  2
  •   hbstha123    4 年前

    使用networkx包解决此问题的另一种方法。

    import networkx as nx
    G = nx.DiGraph()
    G.add_edges_from((r.From, r.To) for r in df.itertuples())
    
    roots = [n for (n, d) in G.in_degree if d == 0]
    print(roots)
    
    leafs = [n for (n, d) in G.out_degree if d == 0]
    print(leafs)
    
    df1 = pd.DataFrame(nx.algorithms.all_simple_paths(G, roots[0], leafs))
    
    for index in df1:
        print (df1.loc[index].to_list())
        
    

    将创建一个有向图,并从添加边 df 节点1是G中唯一的节点,其中in_degree等于0,即它是整个图的根。因此 roots 等于节点0。

    leafs 表示out_degree等于0的有向图上的节点。即[“Node8”、“Node9”、“Node 10”、“节点11”]。

    nx.algorithms.all_simple_paths(G, roots[0], leafs) 提供从节点1开始并在叶中的每个节点处结束的所有路径。

    数据帧的每一行 df1 使用for循环语句打印为列表。

        3
  •  2
  •   Jayyu    4 年前

    我认为@I'mahdi的第一段代码假设一个子节点不会有多个父节点,即图是一个二叉树。当图形不是二叉树时,可以将代码修改为:

    import pandas as pd
    
    def prntPath(lst, node, df, lst_vst):
        for val in df.values:
            if val[0] == node:
                lst.append(val[1])
                # recursely find child node first until no child nodes are found
                prntPath(lst, val[1], df, lst_vst)
                
                # recursion is over
        
        # when no child nodes are found
        if len(lst)>=2:
            if (not lst[-1] in lst_vst) or (not lst[-2] in lst_vst):
                print('-'.join(lst))
        for l in lst:
            lst_vst.add(l)
        lst.pop()
    
    # Two links are added to the graph: 1) from node 4 to node 9. 2) from node 1 to node 11.
     
    df = pd.DataFrame({
              'From':['Node1','Node1', 'Node2', 'Node2', 'Node3', 'Node3', 'Node4', 'Node4','Node5', 'Node6', 'Node7','Node1'],
              'TO'  :['Node2','Node3', 'Node4', 'Node5', 'Node6', 'Node7', 'Node8', 'Node9','Node9', 'Node10', 'Node11','Node12']
            })
    lst_vst = set()
    prntPath(['Node1'],'Node1', df, lst_vst)
    

    结果是:

    Node1-Node2-Node4-Node8
    Node1-Node2-Node4-Node9
    Node1-Node2-Node5-Node9
    Node1-Node3-Node6-Node10
    Node1-Node3-Node7-Node11
    Node1-Node12