繁体   English   中英

从字典导入为多索引pd.DataFrame

[英]Import from a dictionary as a multi-index pd.DataFrame

我有一本字典,它要求像下面这样的多索引:

dict = {'Main1' : {'A1' : {'a1' : 0}, 
                   'A2' : {'a2' : 15}, 
                   'A3' : {'a3' : 22}, 
                   'A4' : {'a4' : 130}},
        'Main2' : {'B1' : {'b1' : 150},
                   'B2' : {'b2' : 30},
                   'B3' : {'b3' : 1}}}

我想将它作为pandas DataFrame导入Python中,如下所示:

col1     col2   col3   col4
Main 1   A1     a1     0
Main 1   A2     a2     15
Main 1   A3     a3     22
Main 1   A4     a4     130
Main 2   B1     b1     150
Main 2   B2     b2     30
Main 2   B3     b3     1

甚至有可能还是我应该尝试寻找另一种导入数据的方法?

您可以这样做:

df = pd.DataFrame([(k1, k2, k3, v) for k1, k23v in dict.items()
                       for k2, k3v in k23v.items()
                       for k3, v in k3v.items()
                       ])
df.columns = ['Col1', 'Col2', 'Col3', 'Col4']

输出:

   Col1 Col2 Col3  Col4
0  Main1  A1  a1    0
1  Main1  A3  a3   22
2  Main1  A2  a2   15
3  Main1  A4  a4  130
4  Main2  B1  b1  150
5  Main2  B2  b2   30
6  Main2  B3  b3    1

这是使用pd.DataFrame.from_dict一种方式:

d = {'Main1' : {'A1' : {'a1' : 0}, 
                'A2' : {'a2' : 15}, 
                'A3' : {'a3' : 22}, 
                'A4' : {'a4' : 130}},
     'Main2' : {'B1' : {'b1' : 150},
                'B2' : {'b2' : 30},
                'B3' : {'b3' : 1}}}

# restructure dictionary to dictionary of tuple keys -> values
d2 = {(i, j, k): d[i][j][k] for i in d.keys()
                            for j in d[i].keys()
                            for k in d[i][j].keys()}

# construct dataframe from dictionary
df = pd.DataFrame.from_dict(d2, orient='index').reset_index()

# split column of tuples to multiple columns
df[['col1', 'col2', 'col3']] = df['index'].apply(pd.Series)

# clean up: remove unwanted columns, rename and sort
df = df.drop('index', 1)\
       .rename(columns={0: 'col4'})\
       .sort_index(axis=1)

print(df)

    col1 col2 col3  col4
0  Main1   A1   a1     0
1  Main1   A2   a2    15
2  Main1   A3   a3    22
3  Main1   A4   a4   130
4  Main2   B1   b1   150
5  Main2   B2   b2    30
6  Main2   B3   b3     1

我发现这样做的另一种方法是使dataframes的字典, concat他们都在一起,然后unstack ,然后删除NaN

dataframes = {k: pd.DataFrame(v) for k,v in d.items()}
dataframe = pd.concat(dataframes, axis=1)
output = dataframe.unstack().dropna()

输出:

Main1  A1  a1      0.0
       A2  a2     15.0
       A3  a3     22.0
       A4  a4    130.0
Main2  B1  b1    150.0
       B2  b2     30.0
       B3  b3      1.0
dtype: float64

暂无
暂无

声明:本站的技术帖子网页,遵循CC BY-SA 4.0协议,如果您需要转载,请注明本站网址或者原文地址。任何问题请咨询:yoyou2525@163.com.

 
粤ICP备18138465号  © 2020-2024 STACKOOM.COM