[英]pandas dataframe groupby columns and aggregate on custom function
[英]Is there a way to write a custom cumulative aggregate function with groupby clause for pandas dataframe?
這是我的 dataframe
+--------+-------------+----------+---------------+------------+-------------+-----------+
| | Customer ID | Quantity | Invoice Value | Date | InvoiceDate | UnitPrice |
+--------+-------------+----------+---------------+------------+-------------+-----------+
| 0 | 500249347 | 0.0 | 0.000 | 2018-01-02 | 2018-01-02 | 0.000 |
+--------+-------------+----------+---------------+------------+-------------+-----------+
| 1 | 500006647 | 1.0 | 33.715 | 2018-01-02 | 2018-01-02 | 33.715 |
+--------+-------------+----------+---------------+------------+-------------+-----------+
| 2 | 500407469 | 1.0 | 33.715 | 2018-01-02 | 2018-01-02 | 33.715 |
+--------+-------------+----------+---------------+------------+-------------+-----------+
| 3 | 500642846 | 0.0 | 0.000 | 2018-01-02 | 2018-01-02 | 0.000 |
+--------+-------------+----------+---------------+------------+-------------+-----------+
| 4 | 500005450 | 1.0 | 33.715 | 2018-01-02 | 2018-01-02 | 33.715 |
+--------+-------------+----------+---------------+------------+-------------+-----------+
| ... | ... | ... | ... | ... | ... | ... |
+--------+-------------+----------+---------------+------------+-------------+-----------+
| 429545 | 500717072 | 1.0 | 45.620 | 2019-03-31 | 2019-03-31 | 45.620 |
+--------+-------------+----------+---------------+------------+-------------+-----------+
| 429546 | 500105174 | 0.0 | 0.000 | 2019-03-31 | 2019-03-31 | 0.000 |
+--------+-------------+----------+---------------+------------+-------------+-----------+
| 429547 | 500069720 | 0.0 | 0.000 | 2019-03-31 | 2019-03-31 | 0.000 |
+--------+-------------+----------+---------------+------------+-------------+-----------+
| 429548 | 500105528 | 0.0 | 0.000 | 2019-03-31 | 2019-03-31 | 0.000 |
+--------+-------------+----------+---------------+------------+-------------+-----------+
| 429549 | 500732322 | 0.0 | 0.000 | 2019-03-31 | 2019-03-31 | 0.000 |
+--------+-------------+----------+---------------+------------+-------------+-----------+
我想提取特征(新列),例如自上次訪問以來每個客戶的天數(每行的快照日期)、上次開票金額、上次非零開票金額、數量和自上次購買以來的天數等。這可以是使用一些自定義累積聚合 function 完成,或者是否有更簡單的方法?
我會建議這樣的事情:
import pandas as pd
df = pd.DataFrame({'customer_id': [13, 16, 13, 13, 16, 16, 13],
'Date': ['2018-01-02', '2019-03-31', '2019-03-31', '2018-01-02', '2018-01-02', '2019-04-31',
'2018-01-02'],
'Invoice_value': [920, 920, 920, 920, 921, 921, 921],
'Unit_price': [1, 2, 3, 4, 6, 7, 8]})
append_data = [df[(df['customer_id'] == ac)].sort_values(by=['Date']).iloc[-1] for ac in df.customer_id.unique()]
自上次訪問以來的時間,我想到了這樣的事情:
df['last_visited']=df.groupby('Customer ID')['Date'].diff()
聲明:本站的技術帖子網頁,遵循CC BY-SA 4.0協議,如果您需要轉載,請注明本站網址或者原文地址。任何問題請咨詢:yoyou2525@163.com.