python从html读取unicode字符

Question

我有此脚本，该脚本从网页读取文本：

page = urllib2.urlopen(url).read()
soup = BeautifulSoup(page);
paragraphs = soup.findAll('p');

for p in paragraphs:
    content = content+p.text+" ";

在网页中，我有以下字符串：

Möddinghofe

我的脚本将其读取为：

M&#195;&#182;ddinghofe

我该如何原样阅读？

Answer 1

希望这对您有帮助

from BeautifulSoup import BeautifulStoneSoup
import cgi

def HTMLEntitiesToUnicode(text):
    """Converts HTML entities to unicode.  For example '&amp;' becomes '&'."""
    text = unicode(BeautifulStoneSoup(text, convertEntities=BeautifulStoneSoup.ALL_ENTITIES))
    return text

def unicodeToHTMLEntities(text):
    """Converts unicode to HTML entities.  For example '&' becomes '&amp;'."""
    text = cgi.escape(text).encode('ascii', 'xmlcharrefreplace')
    return text

text = "&amp;, &reg;, &lt;, &gt;, &cent;, &pound;, &yen;, &euro;, &sect;, &copy;"

uni = HTMLEntitiesToUnicode(text)
htmlent = unicodeToHTMLEntities(uni)

print uni
print htmlent
# &, ®, <, >, ¢, £, ¥, €, §, ©
# &amp;, &#174;, &lt;, &gt;, &#162;, &#163;, &#165;, &#8364;, &#167;, &#169;

参考：将HTML实体转换为Unicode，反之亦然

Answer 2

我建议您看一看BeautifulSoup文档的编码部分。

python从html读取unicode字符

问题描述

2 个解决方案

解决方案1
1 2012-05-14 18:16:22

解决方案2
0 2012-05-14 18:25:47

python从html读取unicode字符

问题描述

2 个解决方案

解决方案1 1 2012-05-14 18:16:22

解决方案2 0 2012-05-14 18:25:47

解决方案1
1 2012-05-14 18:16:22

解决方案2
0 2012-05-14 18:25:47