代码之家  ›  专栏  ›  技术社区  ›  Luca

查找列值低于某个值的时间戳之间的最早时间

  •  0
  • Luca  · 技术社区  · 4 年前

    timestamp,y
    2019-08-01 00:00:00,872.0
    2019-08-01 00:15:00,668.0
    2019-08-01 00:30:00,604.0
    2019-08-01 00:45:00,788.0
    2019-08-01 01:00:00,608.0
    2019-08-01 01:15:00,692.0
    2019-08-01 01:30:00,716.0
    2019-08-01 01:45:00,692.0
    2019-08-01 02:00:00,672.0
    2019-08-01 02:15:00,636.0
    2019-08-01 02:30:00,596.0
    2019-08-01 02:45:00,748.0
    ...
    

    我想做的是,对于这个数据帧中的每个时间段,即下午6点到上午5点之间,我想知道 y

    我正在考虑执行以下伪代码:

    timestamps = list()
    for _, row in df.iterrows():
        found = False
        current = row['timestamp']
        val = row['y']
        if current is between 6 PM and 5 AM:
            if not found and value < threshold:
                found = True
                timestamps.append(current)  
    
    

    但这看起来相当丑陋,而且容易出错,我想知道是否有一种更简洁的pandaish方法来做到这一点?

    1 回复  |  直到 4 年前
        1
  •  3
  •   BeRT2me    4 年前

    设置为使用 df.between_time ,将其设置为DatetimeIndex,并添加任何您想要的可选过滤器:

    df.timestamp = pd.to_datetime(df.timestamp)
    df = df.set_index('timestamp')
    
    threshold = 700
    out = df.between_time('01:00', '02:00')[lambda x: x.y < threshold]
    print(out)
    

                             y
    timestamp
    2019-08-01 01:00:00  608.0
    2019-08-01 01:15:00  692.0
    2019-08-01 01:45:00  692.0
    2019-08-01 02:00:00  672.0
    

    out.resample('d').first()
    
    # Output:
                    y
    timestamp
    2019-08-01  608.0
    

    df.timestamp = pd.to_datetime(df.timestamp)
    df = (df.set_index('timestamp')
            .between_time('18:00', '05:00')
            [lambda x: x.y < threshold]
            .resample('d')
            .first())