![](/img/trans.png)
[英]How to join 2 tables with getting all rows from left table and only matching ones in right
[英]Python pandas get rows from left table and from right table missing in left table
我有左右表,我需要以这种方式合并两个表的FileStamp值:取左表和左表中缺少的右表中的所有值,并按'date'联接:
import pandas as pd
left = pd.DataFrame({'FileStamp': ['T101', 'T102', 'T103', 'T104'], 'date': [20180101, 20180102, 20180103, 20180104]})
right = pd.DataFrame({'FileStamp': ['T501', 'T502'], 'date': [20180104, 20180105]})
就像是
result = pd.merge(left, right, how='outer', on='date')
但是“外面”不是一个好主意。
所需的输出应如下所示
FileStamp_x date FileStamp_y
0 T101 20180101 NaN
1 T102 20180102 NaN
2 T103 20180103 NaN
3 T104 20180104 NaN
4 NaN 20180105 T502
有什么简单的方法可以实现所需的输出吗?
在merge
之前使用isin
进行过滤:
r = right[~right['date'].isin(left['date'])]
print (r)
FileStamp date
1 T502 20180105
result = pd.merge(left, r, how='outer', on='date')
print (result)
FileStamp_x date FileStamp_y
0 T101 20180101 NaN
1 T102 20180102 NaN
2 T103 20180103 NaN
3 T104 20180104 NaN
4 NaN 20180105 T502
您可以在merge
后调整值:
result = pd.merge(left, right, how='outer', on='date')
result['FileStamp_y'] = np.where(result['FileStamp_x'].isnull(), result['FileStamp_y'], np.nan)
结果:
FileStamp_x date FileStamp_y
0 T101 20180101 NaN
1 T102 20180102 NaN
2 T103 20180103 NaN
3 T104 20180104 NaN
4 NaN 20180105 T502
声明:本站的技术帖子网页,遵循CC BY-SA 4.0协议,如果您需要转载,请注明本站网址或者原文地址。任何问题请咨询:yoyou2525@163.com.