我正在尝试基于另一列值替换pandas数据帧中的一个值。
我在下面制作了示例代码来复制这个问题,但本质上我想在现有的数据帧中添加一列,然后根据另一列的值替换占位符信息。我使用的数据帧(不在示例中)基于一个excel文档,该文档将用于根据另一列的值对我希望在新列中返回的信息进行Web抓取。只是想在人们问为什么不开始
ex_list
数据帧中的数据。此外I
只有
希望它替换满足条件的位置,而不是用设置值替换整列。
示例代码
## this would be the excel document df
sample_df = pd.DataFrame({"a":[1,2,3,4,5]})
sample_df["b"] = ""
## this data would be webscrapped using information above
ex_list = [[1, "CHANGE"],[4, "CHANGE"]]
for sub in ex_list:
location = sample_df.loc[sample_df['a']==sub[0], 'b'].iloc[0]
sample_df.replace(location, sub[1])
sample.head()
我也试过这只是一个快速的游戏,但它产生了相同的输出
sample_df = pd.DataFrame({"a":[1,2,3,4,5]})
sample_df["b"] = ""
ex_list = [[1, "CHANGE"],[4, "CHANGE"]]
for sub in ex_list:
sample_df[sample_df['a']==sub[0], 'b'].iloc[0] += sub[1]
sample_df.head()
两个输出相同且没有变化:
a b
0 1
1 2
2 3
3 4
4 5
我希望的结果是
a b
0 1 CHANGE
1 2
2 3
3 4 CHANGE
4 5
如果能再多看一眼,我将不胜感激。我的“定位值”方法是否逻辑错误?我认为.loc/.iloc是最好的,但也许另一种索引方式是最好的?我愿意接受任何解决方案!