比較具有不同鍵的嵌套字典

Question

我試圖將來自 2 個不同來源（因此是兩個字典）的某些值相互比較，以了解哪些值實際上屬於一起。 為了說明，我的兩個字典的較短版本帶有虛擬數據（為清楚起見添加了輸入）

dict_1 = 
{'ins1': {'Start': 100, 'End': 110, 'Size': 10}, 
'ins2': {'Start': 150, 'End': 250, 'Size': 100}, 
'del1': {'Start': 210, 'End': 220, 'Size': 10}, 
'del2': {'Start': 260, 'End': 360, 'Size': 100}, 
'dup1': {'Start': 340, 'End': 350, 'Size': 10, 'Duplications': 3}, 
'dup2': {'Start': 370, 'End': 470, 'Size': 100, 'Duplications': 3}}

dict_2 = 
{'0': {'Start': 100, 'Read': 28, 'Prec': 'PRECISE', 'Size': 10, 'End': 110}, 
'1': {'Start': 500, 'Read': 38, 'Prec': 'PRECISE', 'Size': 100, 'End': 600}, 
'2': {'Start': 210, 'Read': 27, 'Prec': 'PRECISE', 'Size': 10, 'End': 220}, 
'3': {'Start': 650, 'Read': 31, 'Prec': 'IMPRECISE', 'Size': 100, 'End': 750}, 
'4': {'Start': 370, 'Read': 31, 'Prec': 'PRECISE', 'Size': 100, 'End': 470}, 
'5': {'Start': 340, 'Read': 31, 'Prec': 'PRECISE', 'Size': 10, 'End': 350}, 
'6': {'Start': 810, 'Read': 36, 'Prec': 'PRECISE', 'Size': 10, 'End': 820}}

我要比較的是“開始”和“結束”值（以及其他但未在此處指定的值）。 如果它們匹配，我想創建一個與此類似的新 dict (dict_3)：

dict_3 = 
{'ins1': {'Start_d1': 100, 'Start_d2': 100, 'dict_2_ID': '0', etc}
{'del1': {'Start_d1': 210, 'Start_d2': 210, 'dict_2_ID': '2', etc}}

ps 我需要 Start_d1 和 Start_d2，因為它們的數量可能略有不同（+-5）。

我已經在堆棧溢出時嘗試了幾個選項，例如：將具有不同鍵的字典連接到 Pandas 數據幀中（我認為這可以工作，但我在數據幀格式方面遇到了很多麻煩）和：比較 Python 中的兩個字典（僅當字典沒有頂層鍵（比如這里的 ins1、ins2 等）

有人可以讓我開始進一步合作嗎？ 我已經嘗試了很多東西，嵌套字典給我找到的所有解決方案都帶來了麻煩。

Answer 1

也許你可以做這樣的事情：

dict_1 = {'ins1': {'Start': 100, 'End': 110, 'Size': 10},
'ins2': {'Start': 150, 'End': 250, 'Size': 100}, 
'del1': {'Start': 210, 'End': 220, 'Size': 10}, 
'del2': {'Start': 260, 'End': 360, 'Size': 100}, 
'dup1': {'Start': 340, 'End': 350, 'Size': 10, 'Duplications': 3}, 
'dup2': {'Start': 370, 'End': 470, 'Size': 100, 'Duplications': 3}}

dict_2 = {'0': {'Start': 100, 'Read': 28, 'Prec': 'PRECISE', 'Size': 10, 'End': 110},
'1': {'Start': 500, 'Read': 38, 'Prec': 'PRECISE', 'Size': 100, 'End': 600}, 
'2': {'Start': 210, 'Read': 27, 'Prec': 'PRECISE', 'Size': 10, 'End': 220}, 
'3': {'Start': 650, 'Read': 31, 'Prec': 'IMPRECISE', 'Size': 100, 'End': 750}, 
'4': {'Start': 370, 'Read': 31, 'Prec': 'PRECISE', 'Size': 100, 'End': 470}, 
'5': {'Start': 340, 'Read': 31, 'Prec': 'PRECISE', 'Size': 10, 'End': 350}, 
'6': {'Start': 810, 'Read': 36, 'Prec': 'PRECISE', 'Size': 10, 'End': 820}}

dict_3 = {}
for d1 in dict_1:
    for d2 in dict_2:
        if dict_1[d1]["Start"] == dict_2[d2]["Start"] and dict_1[d1]["End"] == dict_2[d2]["End"]:
            dict_3[d1] = {"Start_d1": dict_1[d1]["Start"], "Start_d2": dict_2[d2]["Start"], "dict_2_ID": d2}

print(dict_3)

上面提到的解決方案是n^2 ，這不是很有效。

但是，為了使其更有效（順序n ），您需要以包含"Start"和"End"值作為鍵的方式轉換dict_2 （例如：'S100E110'）然后查找將是恆定時間（字典查找） ref 。 然后，您將能夠執行以下操作：

if str("S"+dict_1[d1]["Start"]+"E"+dict_1[d1]["End"]) in dict_2:    
   # add to dict_3

Answer 2

你可以使用熊貓； 這是一個演示：

import pandas as pd

df1 = pd.DataFrame.from_dict(dict_1, orient='index')
df2 = pd.DataFrame.from_dict(dict_2, orient='index')

res = pd.merge(df1, df2, on=['Start', 'End', 'Size'])

print(res)

   Start  End  Size  Duplications  Read     Prec
0    210  220    10           NaN    27  PRECISE
1    340  350    10           3.0    31  PRECISE
2    370  470   100           3.0    31  PRECISE
3    100  110    10           NaN    28  PRECISE

比較具有不同鍵的嵌套字典

問題描述

2 個解決方案

解決方案1
1 2018-10-03 09:32:44

解決方案2
1 已采納 2018-10-03 10:13:32

比較具有不同鍵的嵌套字典

問題描述

2 個解決方案

解決方案1 1 2018-10-03 09:32:44

解決方案2 1 已采納 2018-10-03 10:13:32

解決方案1
1 2018-10-03 09:32:44

解決方案2
1 已采納 2018-10-03 10:13:32