代码之家  ›  专栏  ›  技术社区  ›  user3768495

在我通过pandas.cut()函数创建了bin之后,如何有效地将每个值标记到bin中?

  •  0
  • user3768495  · 技术社区  · 6 年前

    假设我在数据帧中有一个列是“user_age”,我通过以下方式创建了“user_age_bin”:

    df['user_age_bin']= pd.cut(df['user_age'], bins=[10, 15, 20, 25,30])
    

    然后,我使用“user_age_bin”特性构建了一个机器学习模型。

    接下来,我得到了一条记录,我需要将其放入我的模型中并进行预测。我不想用 user_age 因为模型使用 user_age_bin .那么,我如何转换a 用户年龄 值(比如28)为 user_age_bin ?我知道我可以创建这样的函数:

    def assign_bin(age):
        if age < 10:
            return '<10'
        elif age< 15:
            return '10-15'
         ... etc. etc.
    

    然后执行以下操作:

    user_age_bin = assign_bin(28)
    

    但这种解决方案一点也不优雅。我想一定有更好的办法,对吧?

    编辑:我更改了代码并添加了显式的bin范围。 编辑2:措辞经过编辑,希望现在问题更清楚了。

    1 回复  |  直到 6 年前
        1
  •  4
  •   user3768495    6 年前

    太长,读不下去了 np.digitize 这是一个很好的解决方案。

    在阅读了这里的所有评论和答案以及更多的谷歌搜索后,我想我得到了一个我非常满意的解决方案。感谢你们所有人!

    安装程序

    import pandas as pd
    import numpy as np
    np.random.seed(42)
    
    bins = [0, 10, 15, 20, 25, 30, np.inf]
    labels = bins[1:]
    ages = list(range(5, 90, 5))
    df = pd.DataFrame({"user_age": ages})
    df["user_age_bin"] = pd.cut(df["user_age"], bins=bins, labels=False)
    
    # sort by age 
    print(df.sort_values('user_age'))
    

    输出 :

     user_age  user_age_bin
    0          5             0
    1         10             0
    2         15             1
    3         20             2
    4         25             3
    5         30             4
    6         35             5
    7         40             5
    8         45             5
    9         50             5
    10        55             5
    11        60             5
    12        65             5
    13        70             5
    14        75             5
    15        80             5
    16        85             5
    

    指定类别 :

    # a new age value
    new_age=30
    
    # use this right=True and '-1' trick to make the bins match
    print(np.digitize(new_age, bins=bins, right=True) -1)
    

    输出 :

    4
    
        2
  •  1
  •   m-dz    6 年前

    这是一种有点丑陋的方法,可以理解双重列表,但似乎能做到。

    设置:

    import pandas as pd
    import numpy as np
    np.random.seed(42)
    
    bins = [10, 15, 20, 25, 30, np.Inf]
    labels = bins[1:]
    ages = np.random.randint(10, 35, 10)
    df = pd.DataFrame({"user_age": ages})
    df["user_age_bin"] = pd.cut(df["user_age"], bins=bins, labels=labels)
    print(df)
    

    外出:

       user_age user_age_bin
    0        16         20.0
    1        29         30.0
    2        24         25.0
    3        20         20.0
    4        17         20.0
    5        30         30.0
    6        16         20.0
    7        28         30.0
    8        32          inf
    9        20         20.0
    

    任务:

    # `new_ages` is what you want to assign labels to, used `ages` for simplicity
    new_ages = ages
    ids = [np.argmax([age <= x for x in labels]) for age in new_ages]
    assigned_labels = [labels[i] for i in ids]
    print(pd.DataFrame({"new_ages": new_ages, "assigned_labels": assigned_labels, "user_age_bin": df["user_age_bin"]}))
    

    外出:

       new_ages  assigned_labels user_age_bin
    0        16             20.0         20.0
    1        29             30.0         30.0
    2        24             25.0         25.0
    3        20             20.0         20.0
    4        17             20.0         20.0
    5        30             30.0         30.0
    6        16             20.0         20.0
    7        28             30.0         30.0
    8        32              inf          inf
    9        20             20.0         20.0
    
        3
  •  0
  •   garciparedes    6 年前

    您可以尝试以下操作:

    bins=[10, 15, 20, 25, 30]
    labels = [f'<{bins[0]}', *(f'{a}-{b}' for a, b in zip(bins[:-1], bins[1:])), f'{bins[-1]}>']
    pd.cut(df['user_age'], bins=bins, labels=labels)
    

    请注意,如果您正在使用 python<3.7 您应该用类似语法的格式替换f-string。

        4
  •  0
  •   bbennett36    6 年前

    您不能将字符串放入模型中,因此您需要创建一个映射并跟踪它,或者创建一个单独的列以供以后使用

    def apply_age_bin_numeric(value):
        if value <= 10:
            return 1
        elif value > 10 and value <= 20:
            return 2
        elif value > 21 and value <= 30:
            return 3  
        etc....  
    
    def apply_age_bin_string(value):
        if value <= 10:
            return '<=10'
        elif value > 10 and value <= 20:
            return '11-20'
        elif value > 21 and value <= 30:
            return '21-30' 
        etc....
    
    df['user_age_bin_numeric']= df['user_age'].apply(apply_age_bin_numeric)
    df['user_age_bin_string']= df['user_age'].apply(apply_age_bin_string)  
    

    对于该模型,您将保留 user_age_bin_numeric 然后下降 user_age_bin_string

    在数据进入模型之前,保存一份包含两个字段的数据副本。这样,如果你想显示二进制字段而不是数字二进制字段,你可以将预测与二进制字段的字符串版本相匹配。