代码之家  ›  专栏  ›  技术社区  ›  Frank

如何修改pyspark使用的一行中的一列值

  •  14
  • Frank  · 技术社区  · 8 年前

    我想在userid=22650984时更新值。如何在pyspark平台中实现?谢谢你的帮助。

    >>>xxDF.select('userid','registration_time').filter('userid="22650984"').show(truncate=False)
    18/04/08 10:57:00 WARN TaskSetManager: Lost task 0.1 in stage 57.0 (TID 874, shopee-hadoop-slave89, executor 9): TaskKilled (killed intentionally)
    18/04/08 10:57:00 WARN TaskSetManager: Lost task 11.1 in stage 57.0 (TID 875, shopee-hadoop-slave97, executor 16): TaskKilled (killed intentionally)
    +--------+----------------------------+
    |userid  |registration_time           |
    +--------+----------------------------+
    |22650984|270972-04-26 13:14:46.345152|
    +--------+----------------------------+
    
    3 回复  |  直到 8 年前
        1
  •  24
  •   pault Tanjin    8 年前

    如果您想修改数据帧的子集并保持其余部分不变,最好的选择是使用 pyspark.sql.functions.when() 使用时 filter 或 pyspark.sql.functions.where() 将删除不满足条件的所有行。

    from pyspark.sql.functions import col, when
    
    valueWhenTrue = None  # for example
    
    df.withColumn(
        "existingColumnToUpdate",
        when(
            col("userid") == 22650984,
            valueWhenTrue
        ).otherwise(col("existingColumnToUpdate"))
    )
    

    When将第一个参数作为布尔条件求值。如果条件为 True ,它将返回第二个参数。您可以将多个 when 报表如所示 this post 还有 this post .或使用 otherwise() 指定条件为 False 。

    在本例中,我正在更新现有列 "existingColumnToUpdate" .当 userid 等于指定的值,我将使用 valueWhenTrue 。否则,我们将保持列中的值不变。

        2
  •  0
  •   Aashish Ranjan    5 年前

    根据筛选器更改Dataframe列的值:

    from pyspark.sql.functions import lit new_df = xxDf.filter(xxDf.userid == "22650984").withColumn('clumn_to update', lit(<update_expression>)

        3
  •  -3
  •   karthikr    8 年前

    您可以使用 withColumn 要实现您的目标:

    new_df = xxDf.filter(xxDf.userid = "22650984").withColumn(xxDf.field_to_update, <update_expression>)
    

    update\u表达式将具有更新逻辑-可以是UDF或派生字段等。。

    推荐文章