代码之家  ›  专栏  ›  技术社区  ›  moshevi

大熊猫按任意时间段统计剩余时间

  •  0
  • moshevi  · 技术社区  · 8 年前

    我想把我的 dataframe 到任意时间段,然后创建列,计算每行的剩余时间,直到下一个n个时间段结束。

    例如,此输入:

    df = pd.DataFrame({'dates':pd.date_range('2017-01-01',periods=8,freq='1d')}).set_index('dates')
    
    # the inclusive ends of the time periods
    rolling_dates = ['2017-01-02', '2017-01-05', '2017-01-07', '2017-01-08']  
    
    periods_offests = [0, 1, 2]  # the remaining time periods columns
    

    将产生以下输出:

    dates         periods_expiry_days_0 periods_expiry_days_1   periods_expiry_days_2
    2017-01-01            1                     4.0                       6.0
    2017-01-02            0                     3.0                       5.0
    2017-01-03            2                     4.0                       5.0
    2017-01-04            1                     3.0                       4.0
    2017-01-05            0                     2.0                       3.0
    2017-01-06            1                     2.0                       Nan
    2017-01-07            0                     1.0                       Nan
    2017-01-08            0                     Nan                       Nan
    
    1 回复  |  直到 8 年前
        1
  •  0
  •   T. Ray    8 年前

    def map_offsets(x):
        '''Calculate the day offsets'''
        days = [x for x in (pd_rolling - x).days if x >= 0]
        days += [np.nan] * (len(pd_rolling) - len(days))
    
        return [days[i] for i in periods_offsets]
    
    df = pd.DataFrame({'dates': pd.date_range('2017-01-01', periods=8, freq='1d')}).set_index('dates')
    
    # the inclusive ends of the time periods
    rolling_dates = ['2017-01-02', '2017-01-05', '2017-01-07', '2017-01-08']
    
    # the remaining time periods columns
    periods_offsets = [0, 1, 2]
    
    # Added: Casting to datetime will make offset calculation easier
    pd_rolling = pd.to_datetime(rolling_dates)
    
    # Create new column names for periods
    fmt = 'periods_expiry_days_{}'
    columns = [fmt.format(x) for x in periods_offsets]
    
    # Subtract index from rolling date values, and add to dataframe
    df[columns] = pd.DataFrame(df.index.map(map_offsets).tolist(), index=df.index)