Python：將列表與TF-IDF一起使用

Question

我有以下一段代碼，當前將“令牌”中的所有單詞與“ df”中的每個文檔進行比較。 有什么辦法可以將預定義的單詞列表與文檔（而不是“令牌”）進行比較。

from sklearn.feature_extraction.text import TfidfVectorizer
tfidf_vectorizer = TfidfVectorizer(norm=None)  

list_contents =[]
for index, row in df.iterrows():
    list_contents.append(' '.join(row.Tokens))

# list_contents = df.Content.values

tfidf_matrix = tfidf_vectorizer.fit_transform(list_contents)
df_tfidf = pd.DataFrame(tfidf_matrix.toarray(),columns= [tfidf_vectorizer.get_feature_names()])
df_tfidf.head(10)

任何幫助表示贊賞。 謝謝！

Answer 1

不確定我是否理解正確，但是如果您想讓Vectorizer考慮固定的單詞列表，則可以使用vocabulary參數。

my_words = ["foo","bar","baz"]

# set the vocabulary parameter with your list of words
tfidf_vectorizer = TfidfVectorizer(
    norm=None,
    vocabulary=my_words)  

list_contents =[]
for index, row in df.iterrows():
    list_contents.append(' '.join(row.Tokens))

# this matrix will have only 3 columns because we have forced
# the vectorizer to use just the words foo bar and baz
# so it'll ignore all other words in the documents.
tfidf_matrix = tfidf_vectorizer.fit_transform(list_contents)

Python：將列表與TF-IDF一起使用

問題描述

1 個解決方案

解決方案1
0 已采納 2018-10-21 01:14:57

Python：將列表與TF-IDF一起使用

問題描述

1 個解決方案

解決方案1 0 已采納 2018-10-21 01:14:57

解決方案1
0 已采納 2018-10-21 01:14:57