代码之家  ›  专栏  ›  技术社区  ›  DeduciveR

替换单击流数据中的源

  •  1
  • DeduciveR  · 技术社区  · 7 年前

    我有一个电子商务网站的点击流数据。一些客户可以选择使用贷款/融资选项购买产品。不幸的是,这在下面的reprex中创建了一个新的转介源,标记为“finance”。它还会创建一个或多个新会话。

    我想用同一用户先前会话的源替换源“finance”。

    在本例中,所有会话的观察结果 4-6871.2 & 4-6871.3 将根据会话设置源“direct” 4-6871.1 ,和 3-6871.1 会有'谷歌'作为每个会话的来源 3-6871.0

    我需要在更大的数据集上执行此操作,因此我需要应用逻辑来查找具有“finance”源的会话,并将“finance”实例替换为用户上一个会话的前一个源。

    reprex数据通过 dput

    structure(list(userId = c("6.154032", "6.154032", "6.154032", 
    "6.154032", "6.154032", "6.154032", "6.154032", "6.154032", "6.154032", 
    "8.154036", "8.154036", "8.154036", "8.154036", "8.154036", "8.154036", 
    "8.154036", "8.154036", "8.154036", "8.154036", "8.154036", "8.154036", 
    "8.154036", "8.154036"), session_Id = c("4-6871.0", "4-6871.0", 
    "4-6871.0", "4-6871.1", "4-6871.1", "4-6871.1", "4-6871.2", "4-6871.2", 
    "4-6871.3", "3-6871.0", "3-6871.0", "3-6871.0", "3-6871.0", "3-6871.0", 
    "3-6871.1", "3-6871.1", "3-6871.1", "3-6871.1", "3-6871.1", "3-6871.1", 
    "3-6871.1", "3-6871.1", "3-6871.1"), timeStamp = structure(c(1540294773, 
    1540294828, 1540294841, 1540307321, 1540307341, 1540307718, 1540308709, 
    1540308749, 1540311289, 1540330293, 1540330309, 1540330475, 1540330541, 
    1540330663, 1540331041, 1540331164, 1540331168, 1540331312, 1540331459, 
    1540331465, 1540331579, 1540331603, 1540331630), class = c("POSIXct", 
    "POSIXt"), tzone = "UTC"), source = c("(direct)", "(direct)", 
    "(direct)", "(direct)", "(direct)", "(direct)", "finance", "finance", 
    "finance", "google", "google", "google", "google", "google", 
    "finance", "finance", "finance", "finance", "finance", "finance", 
    "finance", "finance", "finance")), class = c("tbl_df", "tbl", 
    "data.frame"), row.names = c(NA, -23L))
    
    1 回复  |  直到 7 年前
        1
  •  1
  •   Julius Vainora    7 年前

    也许您的完整数据结构会使此解决方案失效,但这里有一个候选方案:

    df <- arrange(df, userId, timeStamp)
    tmp <- rle(df$source)
    tmp$values[tmp$values == "finance"] <- lag(tmp$values)[tmp$values == "finance"]
    df$source <- inverse.rle(tmp)
    table(df$source)
    # (direct)   google 
    #        9       14 
    

    在第一行,我确保顺序正确。然后,假设没有用户的第一个来源可以立即是“finance”,在下面两行中,我用前面的条目替换所有“finance”条目。

    推荐文章