代码之家  ›  专栏  ›  技术社区  ›  Outcast

更改pandas groupby使用的函数值

  •  3
  • Outcast  · 技术社区  · 7 年前

    我正在执行以下操作:

    def percentage(x):
        return x[(x<=5)].count() / x.count() * 100
    
    full_data = full_data.groupby(['Id', 'Week_id'], as_index=False).agg({'Volume': percentage})
    

    但我想这样做 groupby 依次具有多个值,例如 x<=7 , x<=9 , x<=11 等在 percentage 功能。

    要做到这一点,最简单的方法是什么,而不是编写多个函数并调用它们?

    所以基本上我想避免这样做:

    def percentage_1(x):
        return x[(x<=5)].count() / x.count() * 100
    
    full_data_1 = full_data.groupby(['Id', 'Week_id'], as_index=False).agg({'Volume': percentage_1})
    
    def percentage_2(x):
        return x[(x<=7)].count() / x.count() * 100
    
    full_data_2 = full_data.groupby(['Id', 'Week_id'], as_index=False).agg({'Volume': percentage_2})
    
    # etc.
    
    2 回复  |  直到 7 年前
        1
  •  2
  •   jezrael    7 年前

    您可以重写您的函数-创建由布尔值掩码填充的新列,然后聚合 mean 最后一个倍数 100 具有 Series.mul :

    n = 3
    
    full_data['new'] = full_data['Volume'] <= n
    full_data = full_data.groupby(['Id', 'Week_id'])['new'].mean().mul(100).reset_index()
    

    带功能的解决方案:

    def per(df, n):
        df['new'] = df['Volume'] <= n
        return df.groupby(['Id', 'Week_id'])['new'].mean().mul(100).reset_index()
    

    编辑:解决方案来源 github :

    full_data = pd.DataFrame({
            'Id':list('XXYYZZXYZX'),
             'Volume':[2,4,8,1,2,5,8,2,6,4],
             'Week_id':list('aaabbbabac')
    })
    
    print (full_data)
    
    val = 5
    def per(c):
        def f1(x):
            return x[(x<=c)].count() / x.count() * 100
        return f1
    
    full_data2 = full_data.groupby(['Id', 'Week_id']).agg({'Volume': per(val)}).reset_index()
    print (full_data2)
      Id Week_id      Volume
    0  X       a   66.666667
    1  X       c  100.000000
    2  Y       a    0.000000
    3  Y       b  100.000000
    4  Z       a    0.000000
    5  Z       b  100.000000
    

    def percentage(x):
        return x[(x<=val)].count() / x.count() * 100
    
    full_data1 = full_data.groupby(['Id', 'Week_id'], as_index=False).agg({'Volume': percentage})
    
    print (full_data1)
      Id Week_id      Volume
    0  X       a   66.666667
    1  X       c  100.000000
    2  Y       a    0.000000
    3  Y       b  100.000000
    4  Z       a    0.000000
    5  Z       b  100.000000
    
        2
  •  1
  •   Outcast    7 年前

    我想出了一个最简洁的方法来解决我的问题:

    def percentage(x):
        global c
        return x[(x<=c)].count() / x.count() * 100
    
    c=5
    full_data_5 = full_data.groupby(['Id', 'Week_id'], as_index=False).agg({'Volume': percentage})
    
    c=7
    full_data_7 = full_data.groupby(['Id', 'Week_id'], as_index=False).agg({'Volume': percentage})
    
    c=9
    full_data_9 = full_data.groupby(['Id', 'Week_id'], as_index=False).agg({'Volume': percentage})
    
    # etc
    

    但是,我使用的是全局变量,这是一个有争议的实践。