代码之家  ›  专栏  ›  技术社区  ›  JohnE

为什么比较顺序对这个应用/λ不等式很重要?

  •  3
  • JohnE  · 技术社区  · 11 年前

    抱歉,这不是一个好标题。不过,简单的例子是:

    (熊猫版本0.16.1)

    df = pd.DataFrame({ 'x':range(1,5), 'y':[1,1,1,9] })
    

    工作正常:

    df.apply( lambda x: x > x.mean() )
    
           x      y
    0  False  False
    1  False  False
    2   True  False
    3   True   True
    

    这不是应该一样吗?

    df.apply( lambda x: x.mean() < x )
    ---------------------------------------------------------------------------
    TypeError                                 Traceback (most recent call last)
    <ipython-input-467-6f32d50055ea> in <module>()
    ----> 1 df.apply( lambda x: x.mean() < x )
    
    C:\Users\ei\AppData\Local\Continuum\Anaconda\lib\site-packages\pandas\core\frame.pyc in apply(self, func, axis, broadcast, raw, reduce, args, **kwds)
       3707                     if reduce is None:
       3708                         reduce = True
    -> 3709                     return self._apply_standard(f, axis, reduce=reduce)
       3710             else:
       3711                 return self._apply_broadcast(f, axis)
    
    C:\Users\ei\AppData\Local\Continuum\Anaconda\lib\site-packages\pandas\core\frame.pyc in _apply_standard(self, func, axis, ignore_failures, reduce)
       3797             try:
       3798                 for i, v in enumerate(series_gen):
    -> 3799                     results[i] = func(v)
       3800                     keys.append(v.name)
       3801             except Exception as e:
    
    <ipython-input-467-6f32d50055ea> in <lambda>(x)
    ----> 1 df.apply( lambda x: x.mean() < x )
    
    C:\Users\ei\AppData\Local\Continuum\Anaconda\lib\site-packages\pandas\core\ops.pyc in wrapper(self, other, axis)
        586             return NotImplemented
        587         elif isinstance(other, (np.ndarray, pd.Index)):
    --> 588             if len(self) != len(other):
        589                 raise ValueError('Lengths must match to compare')
        590             return self._constructor(na_op(self.values, np.asarray(other)),
    
    TypeError: ('len() of unsized object', u'occurred at index x')
    

    举个反例,这两种方法都有效:

    df.mean() < df
    
    df > df.mean()
    
    2 回复  |  直到 11 年前
        1
  •  3
  •   Anand S Kumar    11 年前

    编辑

    终于找到了这个bug- Issue 9369

    如问题所示-

    左=0>s工作(例如python标量)。所以我认为这是 被视为一个0维数组(它是一个np.int64)(当 调用。)我会标记为bug。请随意挖掘

    将比较运算符与 numpy 比较运算符左侧的数据类型(如np.int64或np.float64等)。一个简单的修复方法可能如@santon在回答中所指出的那样,将数字转换为python标量,而不是使用 numpy的复数 标量。


    旧版本:

    我试过潘达斯0.16.2。

    我对你的原始df做了以下操作-

    In [22]: df['z'] = df['x'].mean() < df['x']
    
    In [23]: df
    Out[23]:
       x  y      z
    0  1  1  False
    1  2  1  False
    2  3  1   True
    3  4  9   True
    
    In [27]: df['z'].mean() < df['z']
    ---------------------------------------------------------------------------
    TypeError                                 Traceback (most recent call last)
    <ipython-input-27-afc8a7b869b4> in <module>()
    ----> 1 df['z'].mean() < df['z']
    
    C:\Anaconda3\lib\site-packages\pandas\core\ops.py in wrapper(self, other, axis)
        586             return NotImplemented
        587         elif isinstance(other, (np.ndarray, pd.Index)):
    --> 588             if len(self) != len(other):
        589                 raise ValueError('Lengths must match to compare')
        590             return self._constructor(na_op(self.values, np.asarray(other)),
    
    TypeError: len() of unsized object
    

    对我来说,这似乎是一个bug,我可以将布尔means与int进行比较,反之亦然,但唯一的问题是将布尔mean与布尔值进行比较(尽管我认为将mean()用于布尔值是不合理的)-

    In [24]: df['z'] < df['x']
    Out[24]:
    0    True
    1    True
    2    True
    3    True
    dtype: bool
    
    In [25]: df['z'] < df['x'].mean()
    Out[25]:
    0    True
    1    True
    2    True
    3    True
    Name: z, dtype: bool
    
    In [26]: df['x'].mean() < df['z']
    Out[26]:
    0    False
    1    False
    2    False
    3    False
    Name: z, dtype: bool
    

    我尝试在Pandas 0.16.1中复制该问题,也可以使用-

    In [10]: df['x'].mean() < df['x']
    ---------------------------------------------------------------------------
    TypeError                                 Traceback (most recent call last)
    <ipython-input-10-4e5dab1545af> in <module>()
    ----> 1 df['x'].mean() < df['x']
    
    /opt/anaconda/envs/np18py27-1.9/lib/python2.7/site-packages/pandas/core/ops.pyc in wrapper(self, other, axis)
        586             return NotImplemented
        587         elif isinstance(other, (np.ndarray, pd.Index)):
    --> 588             if len(self) != len(other):
        589                 raise ValueError('Lengths must match to compare')
        590             return self._constructor(na_op(self.values, np.asarray(other)),
    
    TypeError: len() of unsized object
    
    In [11]: df['x'] < df['x'].mean()
    Out[11]: 
    0     True
    1     True
    2    False
    3    False
    Name: x, dtype: bool
    

    这似乎也是Pandas版本0.16.2中修复的一个bug(布尔值与整数混合时除外)。我建议使用-

    pip install pandas --upgrade
    

        2
  •  2
  •   santon    11 年前

    我认为这与大于运算符的重载有关。当使用重载函数时,如果左侧或右侧的数据类型不同,则顺序很重要。(Python有一种复杂的方法来确定使用哪个重载函数。) mean() (即 numpy.float64 )到一个简单的浮点数:

    df.apply( lambda x: float(x.mean()) < x )
    

    出于某种原因,熊猫守则似乎正在处理 浮点数64 这可能是它失败的原因。