主题建模错误（doc2bow 需要输入 unicode 令牌数组，而不是单个字符串）

Question

from nltk.tokenize import RegexpTokenizer
#from stop_words import get_stop_words
from gensim import corpora, models 
import gensim
import os
from os import path
from time import sleep

filename_2 = "buisness1.txt"
file1 = open(filename_2, encoding='utf-8')  
Reader = file1.read()
tdm = []

# Tokenized the text to individual terms and created the stop list
tokens = Reader.split()
#insert stopwords files
stopwordfile = open("StopWords.txt", encoding='utf-8')  

# Use this to read file content as a stream  
readstopword = stopwordfile.read() 
stop_words = readstopword.split() 

for r in tokens:  
    if not r in stop_words: 
        #stopped_tokens = [i for i in tokens if not i in en_stop]
        tdm.append(r)

dictionary = corpora.Dictionary(tdm)
corpus = [dictionary.doc2bow(i) for i in tdm]
sleep(3)
#Implemented the LdaModel
ldamodel = gensim.models.ldamodel.LdaModel(corpus, num_topics=10, id2word = dictionary)
print(ldamodel.print_topics(num_topics=1, num_words=1))

我正在尝试使用包含停用词的单独 txt 文件删除停用词。 在我删除停用词后，我将附加停用词中不存在的文本文件的单词。 我收到错误doc2bow expects an array of unicode tokens on input, not a single string dictionary = corpora.Dictionary(tdm)处的单个字符串。

谁能帮我更正我的代码

Answer 1

这几乎可以肯定是重复的，但请改用它：

dictionary = corpora.Dictionary([tdm])

主题建模错误（doc2bow 需要输入 unicode 令牌数组，而不是单个字符串）

问题描述

1 个解决方案

解决方案1
0 已采纳 2021-04-29 09:35:54

主题建模错误（doc2bow 需要输入 unicode 令牌数组，而不是单个字符串）

问题描述

1 个解决方案

解决方案1 0 已采纳 2021-04-29 09:35:54

解决方案1
0 已采纳 2021-04-29 09:35:54