簡體   English   中英

從分組的pandas數據框中繪制堆積圖

[英]Plotting stacked plot from grouped pandas data frame

我有一個數據框,如下所示。首先,我想要計算每個日期的每個狀態。 例如2017-11-02中“完成”的數量是2.我想要一個這樣的疊加圖。

                   status              start_time                end_time  \
0             COMPLETED 2017-11-01 19:58:54.726 2017-11-01 20:01:05.414   
1             COMPLETED 2017-11-02 19:43:04.000 2017-11-02 19:47:54.877   
2     ABANDONED_BY_USER 2017-11-03 23:36:19.059 2017-11-03 23:36:41.045   
3  ABANDONED_BY_TIMEOUT 2017-10-31 17:02:38.689 2017-10-31 17:12:38.844   
4             COMPLETED 2017-11-02 19:35:33.192 2017-11-02 19:42:51.074   

這是數據幀的csv:

status,start_time,end_time
COMPLETED,2017-11-01 19:58:54.726,2017-11-01 20:01:05.414
COMPLETED,2017-11-02 19:43:04.000,2017-11-02 19:47:54.877
ABANDONED_BY_USER,2017-11-03 23:36:19.059,2017-11-03 23:36:41.045
ABANDONED_BY_TIMEOUT,2017-10-31 17:02:38.689,2017-10-31 17:12:38.844
COMPLETED,2017-11-02 19:35:33.192,2017-11-02 19:42:51.074
ABANDONED_BY_TIMEOUT,2017-11-02 19:35:33.192,2017-11-02 19:42:51.074

為達到這個:

df_['status'].astype('category')
df_ = df_.set_index('start_time')
grouped = df_.groupby('status')
color = {'COMPLETED':'green','ABANDONED_BY_TIMEOUT':'blue',"MISSED":'red',"ABANDONED_BY_USER":'yellow'}

for key_, group in grouped:
   print(key_)
   df_ = group.groupby(lambda x: x.date).count()
   print(df_)
   df_['status'].plot(label=key_,kind='bar',stacked=True,\
   color=color[key_],rot=90)
plt.show()

以下輸出是:

ABANDONED_BY_TIMEOUT
            status  end_time  
2017-10-31       1         1       
ABANDONED_BY_USER
            status  end_time  
2017-11-03       1         1            
COMPLETED
            status  end_time  
2017-11-01       1         1             
2017-11-02       2         2 

從上面的代碼繪制

我們可以看到這里的問題僅考慮過去兩個日期'2017-11-01'和'2017-11-02'而不是所有類別中的所有日期。 我怎樣才能解決這個問題呢?我歡迎采用全新的堆積方法。謝謝。

import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns

df_ = pd.read_csv('sam.csv')
df_['date'] = pd.to_datetime(df_['start_time']).dt.date
df_ = df_.set_index('start_time')


grouped = pd.DataFrame(df_.groupby(['date', 'status']).size().reset_index(name="count")).pivot(columns='status', index='date', values='count')
print(grouped)
sns.set()

grouped.plot(kind='bar', stacked=True)

# g = grouped.plot(x='date', kind='bar', stacked=True)
plt.show()

輸出:

在此輸入圖像描述

嘗試使用pandas.crosstab重構df_

color = ['blue', 'yellow', 'green', 'red']
df_xtab = pd.crosstab(df_.start_time.dt.date, df_.status)

DataFrame將如下所示:

status      ABANDONED_BY_TIMEOUT  ABANDONED_BY_USER  COMPLETED
start_time                                                    
2017-10-31                     1                  0          0
2017-11-01                     0                  0          1
2017-11-02                     1                  0          2
2017-11-03                     0                  1          0

並且將更容易繪圖。

df_xtab.plot(kind='bar',stacked=True, color=color, rot=90)

在此輸入圖像描述

使用seaborn library barplot及其色調

碼:

import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns

df_ = pd.read_csv('sam.csv')
df_['date'] = pd.to_datetime(df_['start_time']).dt.date
df_ = df_.set_index('start_time')

print(df_)

grouped = pd.DataFrame(df_.groupby(['date', 'status']).size().reset_index(name="count"))
print(grouped)

g = sns.barplot(x='date', y='count', hue='status', data=grouped)
plt.show()

輸出: 在此輸入圖像描述


數據:

status,start_time,end_time
COMPLETED,2017-11-01 19:58:54.726,2017-11-01 20:01:05.414
COMPLETED,2017-11-02 19:43:04.000,2017-11-02 19:47:54.877
ABANDONED_BY_USER,2017-11-03 23:36:19.059,2017-11-03 23:36:41.045
ABANDONED_BY_TIMEOUT,2017-10-31 17:02:38.689,2017-10-31 17:12:38.844
COMPLETED,2017-11-02 19:35:33.192,2017-11-02 19:42:51.074
ABANDONED_BY_TIMEOUT,2017-11-02 19:35:33.192,2017-11-02 19:42:51.074

在此輸入圖像描述

暫無
暫無

聲明:本站的技術帖子網頁,遵循CC BY-SA 4.0協議,如果您需要轉載,請注明本站網址或者原文地址。任何問題請咨詢:yoyou2525@163.com.

 
粵ICP備18138465號  © 2020-2024 STACKOOM.COM