代码之家  ›  专栏  ›  技术社区  ›  nasa313

如何对多个panda列进行t检验

  •  1
  • nasa313  · 技术社区  · 2 年前

    我想写一个在上运行t检验的代码(只有几行) Product 和 Purchase_cost , warranty_years 和 service_cost 同时

    # dataset 
    
    import pandas as pd
    from scipy.stats import ttest_ind
    
    data = {'Product': ['laptop', 'printer','printer','printer','laptop','printer','laptop','laptop','printer','printer'],
            'Purchase_cost': [120.09, 150.45, 300.12, 450.11, 200.55,175.89,124.12,113.12,143.33,375.65],
            'Warranty_years':[3,2,2,1,4,1,2,3,1,2],
            'service_cost': [5,5,10,4,7,10,4,6,12,3]
        
            }
    
    df = pd.DataFrame(data)
    
    print(df)
    
    

    的代码尝试 产品 & 采购成本 。我想进行t检验 产品 & 保修_年 和 产品 & service cost

    
    #define samples
    group1 = df[df['Product']=='laptop']
    group2 = df[df['Product']=='printer']
    
    #perform independent two sample t-test
    ttest_ind(group1['Purchase_cost'], group2['Purchase_cost'])
    
    
    1 回复  |  直到 2 年前
        1
  •  1
  •   mozway    2 年前

    ttest_ind 可以处理2D(ND)输入:

    cols = df.columns.difference(['Product'])
    # or with an explicit list
    # cols = ['Purchase_cost', 'Warranty_years', 'service_cost']
    
    group1 = df[df['Product']=='laptop']
    group2 = df[df['Product']=='printer']
    out = pd.DataFrame(ttest_ind(group1[cols], group2[cols]),
                       columns=cols, index=['statistic', 'pvalue'])
    

    如果不是这样,你可以使用字典理解在你的列上循环:

    out = pd.DataFrame({c: ttest_ind(group1[c], group2[c]) for c in cols},
                        index=['statistic', 'pvalue'])
    

    输出:

               Purchase_cost  Warranty_years  service_cost
    statistic      -1.861113        3.513240     -0.919464
    pvalue          0.099760        0.007924      0.384738
    

    推广到更多对

    如果您的产品不仅仅是笔记本电脑/打印机,并且希望比较所有配对,您可以概括为:

    from itertools import combinations
    
    cols = df.columns.difference(['Product'])
    
    g = df.groupby('Product')[cols]
    
    out = pd.concat({(a,b): pd.DataFrame(ttest_ind(g.get_group(a), g.get_group(b)),
                                         columns=cols, index=['statistic', 'pvalue'])
                     for a, b in combinations(df['Product'].unique(), 2)
                    }, names=['product1', 'product2'])
    

    带有额外类别(电话)的示例输出:

                                 Purchase_cost  Warranty_years  service_cost
    product1 product2                                                       
    laptop   printer  statistic      -1.861113        3.513240     -0.919464
                      pvalue          0.099760        0.007924      0.384738
             phone    statistic      -1.945836        2.988072      2.766417
                      pvalue          0.109251        0.030515      0.039533
    printer  phone    statistic      -1.286968        0.423659      1.893370
                      pvalue          0.239026        0.684528      0.100178
    

    如果您有许多组合,请注意,您可能应该对数据进行后处理以说明 multiple testing 。