Pandas组连续和label长度

Question

I want get consecutive length labeled data我想获得连续长度标记的数据

I want:我想：

then I can calculate the mean of "b" column by group "c".然后我可以按组“c”计算“b”列的平均值。 tried with shift and cumsum and cumcount all not work.尝试使用 shift 和 cumsum 和 cumcount 都不起作用。

Answer 1

Use GroupBy.transform by consecutive groups and then set 0 if not 1 in a column:按连续组使用GroupBy.transform ，然后a列中设置0如果不是1 ：

df['c1'] = (df.groupby(df.a.ne(df.a.shift()).cumsum())['a']
              .transform('size')
              .where(df.a.eq(1), 0))
print (df)
    a  b  c  c1
0   1  1  1   1
1   0  2  0   0
2   1  3  2   2
3   1  2  2   2
4   0  1  0   0
5   1  3  3   3
6   1  1  3   3
7   1  3  3   3
8   0  2  0   0
9   1  2  2   2
10  1  1  2   2

If there are only 0, 1 values is possible multiple by a :如果只有0, 1值可能是a的倍数：

df['c1'] = (df.groupby(df.a.ne(df.a.shift()).cumsum())['a']
              .transform('size')
              .mul(df.a))
print (df)
    a  b  c  c1
0   1  1  1   1
1   0  2  0   0
2   1  3  2   2
3   1  2  2   2
4   0  1  0   0
5   1  3  3   3
6   1  1  3   3
7   1  3  3   3
8   0  2  0   0
9   1  2  2   2
10  1  1  2   2

Pandas组连续和label长度

问题描述

1 个解决方案

解决方案1
0 2022-08-17 05:49:40

Pandas组连续和label长度

问题描述

1 个解决方案

解决方案1 0 2022-08-17 05:49:40

解决方案1
0 2022-08-17 05:49:40