如何从pandas中的另一列中减去字符串类型列值

Question

我有这样的数据框

df
col1         col2          col3
 A        black berry      black
 B        green apple      green
 C        red wine          red

我想从col2值中减去col3值，结果看起来像

df1
col1        col2        col3
  A         berry       black
  B         apple       green
  C          wine        red

如何使用pandas以有效的方式做到这一点

Answer 1

使用list comprehension with replace和split ：

df['col2'] = [a.replace(b, '').strip() for a, b in zip(df['col2'], df['col3'])]
print (df)
  col1   col2   col3
0    A  berry  black
1    B  apple  green
2    C   wine    red

如果顺序不重要，请将拆分值转换为set和subtract：

df['col2'] = [' '.join(set(a.split())-set([b])) for a, b in zip(df['col2'], df['col3'])]
print (df)
  col1   col2   col3
0    A  berry  black
1    B  apple  green
2    C   wine    red

或者使用if条件和join生成器：

df['col2'] = [' '.join(c for c in a.split() if c != b) for a, b in zip(df['col2'], df['col3'])]

表现：

这是用于生成上面的perfplot的设置：

def calculation(val):
    return val[0].replace(val[1],'').strip()


def regex(df):
    df.col2=df.col2.replace(regex=r'(?i)'+ df.col3,value="")
    return df

def lambda_f(df):
    df["col2"] = df.apply(lambda x: x["col2"].replace(x["col3"], "").strip(), axis=1)
    return df

def apply(df):
    df['col2'] = df[['col2','col3']].apply(calculation, axis=1)
    return df

def list_comp1(df):
    df['col2'] = [a.replace(b, '').strip() for a, b in zip(df['col2'], df['col3'])]
    return df

def list_comp2(df):
    df['col2'] = [' '.join(set(a.split())-set([b])) for a, b in zip(df['col2'], df['col3'])]
    return df

def list_comp3(df):
    df['col2'] = [' '.join(c for c in a.split() if c != b) for a, b in zip(df['col2'], df['col3'])]
    return df


def make_df(n):
    d = {'col1': {0: 'A', 1: 'B', 2: 'C'}, 'col2': {0: 'black berry', 1: 'green apple', 2: 'red wine'}, 'col3': {0: 'black', 1: 'green', 2: 'red'}}
    df = pd.DataFrame(d)
    df = pd.concat([df] * n * 100, ignore_index=True)
    return df

perfplot.show(
    setup=make_df,
    kernels=[regex, lambda_f, apply, list_comp1,list_comp2,list_comp3],
    n_range=[2**k for k in range(2, 10)],
    logx=True,
    logy=True,
    equality_check=False,  # rows may appear in different order
    xlabel='len(df)')

Answer 2

单线解决方案：

df["col2"] = df.apply(lambda x: x["col2"].replace(x["col3"], "").strip(), axis=1)

Answer 3

不需要for loop replace ，请注意这是由row替换

df.col2=df.col2.replace(regex=r'(?i)'+ df.col3,value="")
df
Out[627]: 
  col1    col2   col3
0    A   berry  black
1    B   apple  green
2    C    wine    red

更多信息

  col1   col2   col3
0    A  berry  black
1    B  apple  green
2    C   wine  apple# different row with row 2 , but same value 

df.col2.replace(regex=r'(?i)'+ df.col3,value="")
Out[629]: 
0    berry
1    apple# would not be replaced 
2     wine

Answer 4

我们可以使用apply方法：

def calculation(val):
    return val[0].replace(val[1],'').strip()

df['col4'] = df[['col2','col3']].apply(calculation, axis=1)

df:
  col1         col2   col3   col4
0    A  black berry  black  berry
1    B  green apple  green  apple
2    C    red wine     red   wine

如何从pandas中的另一列中减去字符串类型列值

问题描述

4 个解决方案

解决方案1
3 已采纳 2019-02-22 14:06:47

解决方案2
2 2019-02-22 14:23:12

解决方案3
2 2019-02-22 14:34:06

解决方案4
1 2019-02-22 14:18:53

如何从pandas中的另一列中减去字符串类型列值

问题描述

4 个解决方案

解决方案1 3 已采纳 2019-02-22 14:06:47

解决方案2 2 2019-02-22 14:23:12

解决方案3 2 2019-02-22 14:34:06

解决方案4 1 2019-02-22 14:18:53

解决方案1
3 已采纳 2019-02-22 14:06:47

解决方案2
2 2019-02-22 14:23:12

解决方案3
2 2019-02-22 14:34:06

解决方案4
1 2019-02-22 14:18:53