这很可能以前有人问过,但我很难阐明我的问题。
在我的数据中,我有3个变量,LOCATION、TOPIC和RESPONSE。我想按位置计算TOPIC和RESPONSE的每个组合的分布。
创建玩具数据并执行初始数据准备
responses <- data.frame(LOCATION = c("LOC_A", "LOC_A", "LOC_A", "LOC_A", "LOC_A",
"LOC_A", "LOC_A", "LOC_A",
"LOC_B", "LOC_B", "LOC_B", "LOC_B", "LOC_B",
"LOC_C", "LOC_C", "LOC_C", "LOC_C", "LOC_C", "LOC_C",
"LOC_C", "LOC_C", "LOC_C", "LOC_C", "LOC_C", "LOC_C",
"LOC_C", "LOC_C", "LOC_C", "LOC_C", "LOC_C"),
TOPIC = c("Dogs", "Dogs", "Dogs", "Dogs", "Dogs", "Dogs",
"Lizards", "Lizards", "Lizards",
"Lizards", "Lizards", "Lizards", "Lizards", "Lizards",
"Lizards", "Lizards", "Snakes", "Snakes", "Snakes", "Snakes", "Snakes",
"Snakes", "Dogs", "Snakes", "Dogs", "Snakes", "Dogs",
"Snakes", "Dogs", "Snakes"),
RESP = c("Agree", "Disagree", "Agree", "Disagree", "Agree",
"Disagree", "Agree", "Disagree",
"Agree", "Disagree", "Agree", "Disagree", "Neither", "Agree",
"Neither", "Agree", "Neither", "Agree", "Neither",
"Agree", "Neither", "Agree", "Agree", "Neither",
"Agree", "Neither", "Agree", "Disagree", "Disagree",
"Neither"))
# Obtain counts for each combination of levels
distribution <- responses %>%
table() %>%
as.data.frame() %>%
# Make it more readable
dplyr::arrange(LOCATION, TOPIC, RESP)
下面是一个使用循环创建所需输出的示例解决方案:
# ugly loop solution :(
# Initialize output container
out <- list()
# Iterate over each location
for(loc in unique(distribution$LOCATION)){
# Subset distribution for this location
thisDist <- dplyr::filter(distribution, LOCATION == loc)
# Calculate percent of each response for this location
thisDist$percent <- thisDist$Freq/sum(thisDist$Freq)
# Store distribution df with percent column
out[[loc]] <- thisDist
}
# combine output into single df
out <- do.call("rbind", out)
我想要的是一个简洁明了的解决方案。下面是一些伪代码,它描述了我想象中的解决方案。
# Imaginary tidyverse solution :)
out <- distribution %>%
group_by(LOCATION, TOPIC, RESP) %>%
summarise(#percent = Freq/(sum(<all-Freq-values-for-this-group's-LOCATION-value>))
)
我在这里要做的是获得当前组的LOCATION值的所有Freq值的总和。有没有一种很好的方法可以在一个小组中做到这一点?
谢谢你的阅读,我希望这不是完全无法理解的。