代码之家  ›  专栏  ›  技术社区  ›  DNB5brims

有没有简单的方法来获取python中的峰值和最低值?[闭门]

  •  1
  • DNB5brims  · 技术社区  · 7 年前

    请考虑以下格式的我的数据:

    20180101,10
    20180102,20
    20180103,15
    ....
    

    第一个是日期,第二个是销售的产品数量,而不是将所有这些数据插入数据库,并使用select max xxxx SQL语句来找出一段时间内的最大数量,是否有任何速记或有用的库可以用于此目的?谢谢

    5 回复  |  直到 7 年前
        1
  •  1
  •   JBB    7 年前

    这可能是一个有偏见的答案,但熊猫真的很适合处理这样的数据。而您可以使用元组、列表等来完成这种操作。 熊猫提供了更多的功能。例如:

    import pandas as pd
    data = [[20180101,15], [20180102,10], [20180103,12],[20180104,10]]
    df = pd.DataFrame(data=data, columns=['date', 'products'])
    # if your data is in csv, excel, database... whatever... you can easily pull
    # df = pd.read_csv('name') || pd.read_excel() || pd.read_sql()
    df
    Out[2]: 
           date  products
    0  20180101        15
    1  20180102        10
    2  20180103        12
    3  20180104        10
    
    # It helps to use datetime format to perform operations on the data
    # Operations make reference to an "index" in the dataframe
    df.index = pd.to_datetime(df['date'], format="%Y%m%d")  #strftime format
    df
    Out[3]: 
                    date  products
    date                          
    2018-01-01  20180101        15
    2018-01-02  20180102        10
    2018-01-03  20180103        12
    2018-01-04  20180104        10
    
    # Now we can drop that date column...
    df.drop(columns='date', inplace=True)
    df
    Out[4]: 
                products
    date                
    2018-01-01        15
    2018-01-02        10
    2018-01-03        12
    2018-01-04        10
    
    # Yes, there are ways to do the above in shorthand... lots of info on pandas on SO
    # I want you to see the individual steps we are taking to keep simple
    
    # Now is when the fun begins
    df.rolling(2).sum()  # prints a rolling 2-day sum
    Out[5]: 
                products
    date                
    2018-01-01       NaN
    2018-01-02      25.0
    2018-01-03      22.0
    2018-01-04      22.0
    
    df.rolling(3).mean()  # prints a rolling 3-day average
    Out[6]: 
                 products
    date                 
    2018-01-01        NaN
    2018-01-02        NaN
    2018-01-03  12.333333
    2018-01-04  10.666667
    
    df.resample('W').sum()  # Resamples the data so you can look on a weekly basis
    Out[7]: 
                products
    date                
    2018-01-07        47
    
    df.rolling(2).max() # max number of products over a rolling two-day period
    Out[9]: 
                products
    date                
    2018-01-01       NaN
    2018-01-02      15.0
    2018-01-03      12.0
    2018-01-04      12.0
    
        2
  •  1
  •   filippo    7 年前

    Pandas 这就是你想要的自由。

    让我举一个例子:

    import numpy as np
    import pandas as pd
    
    # let's build a dummy dataset
    index = pd.date_range(start="1/1/2015", end="31/12/2018")
    df = pd.DataFrame(np.random.randint(100, size=len(index)),
                      columns=["sales"], index=index)
    
    >>> df.head()
                sales
    2015-01-01     32
    2015-01-02      0
    2015-01-03     12
    2015-01-04     77
    2015-01-05     86
    

    现在,假设您希望每月汇总销售额:

    >>> df["sales"].groupby(pd.Grouper(freq="1M")).sum()
    
    2015-01-31    1441
    2015-02-28    1164
    2015-03-31    1624
    2015-04-30    1629
    2015-05-31    1427
    [...]
    

    还是以学期为基础

    df["sales"].groupby(pd.Grouper(freq="6M", closed="left", label="right")).sum()    
    2015-06-30    8921
    2015-12-31    9365
    2016-06-30    9820
    2016-12-31    8881
    2017-06-30    8773
    2017-12-31    8709
    2018-06-30    9481
    2018-12-31    9522
    2019-06-30      51
    

    出于某种原因 Grouper binning with six months freq与31/12销售有一些问题,它在2019年将其放入一个新的bin,如果我发现任何问题,查看它会让你知道。。。或者如果其他人想发表评论,请发表评论

    或者你想知道哪个学期最好:

    >>> df["sales"].groupby(pd.Grouper(freq="6M")).sum().idxmax()              
    Timestamp('2016-06-30 00:00:00', freq='6M')
    
        3
  •  0
  •   Steven G    7 年前

    你应该使用 pandas

    假设您的日期列名为“date”,并且是datetime数据类型:

    import pandas as pd
    df = pd.DataFrame(data)
    df = df.set_index('date')
    df.groupby(pd.Grouper(freq='1M')).max()
    

    会给你每个月的最大频率可以改变到任何你喜欢的频率。

        4
  •  0
  •   Alex_P    7 年前

    我试过@Patrick Artner的评论:

    a = (20180101,10)
    b = (20180102,20)
    c = (20180103,15)
    d = (a,b,c)
    maximum = max( d, key = lambda x:x[1])
    minimum = min(d, key= lambda x:x[1])
    print(minimum)
    

    也许这会给我们一些启发。

        5
  •  -1
  •   Sreeragh A R    7 年前

    如果这是期望的结果,请回答。

    data = [{'date':1, 'products_sold': 2}, {'date':2, 'products_sold': 5},{'date':5, 'products_sold': 2}]
    start_date = 1
    end_date = 2
    max_value_in_period = max(x['products_sold'] for x in data if x['date'] >= start_date and x['date'] <= end_date)
    print(max_value_in_period)