使用BeautifulSoup提取部分字符串

Question

我正在使用Beautiful Soup 4解析html文檔並提取數據。

我想從此標簽獲取時間值：

<span style="font-size:9.0pt;font-family:Arial;color:#666666"> 20 min <b>Start time: </b> 10 min <b>Other time: </b> 0 min</span>

IE：20分鍾，10分鍾

Answer 1

這有幫助嗎？

from BeautifulSoup import BeautifulSoup
from BeautifulSoup import Tag

soup = BeautifulSoup("<span style=\"font-size:9.0pt;font-family:Arial;color:#666666\"> 20 min <b>Start time: </b> 10 min <b>Other time: </b> 0 min</span>")
span = soup.find('span')
for e in span.contents:
 if type(e) is Tag:
   print "found a tag:", e.name
 else:
   print "found text:", e

輸出：

found text:  20 min
found a tag: b
found text:  10 min
found a tag: b
found text:  0 min

Answer 2

這是應該做的：

from bs4 import BeautifulSoup

ss = """<span style="font-size:9.0pt;font-family:Arial;color:#666666"> 20 min <b>Start time:     </b> 10 min <b>Other time: </b> 0 min</span>"""
soup = BeautifulSoup(ss)
timetext = soup.span.text
start_time = timetext.split("Start time:")[1].split("min")[0].strip()

我已經摘錄了other_time作為練習！

使用BeautifulSoup提取部分字符串

問題描述

2 個解決方案

解決方案1
2 已采納 2014-12-12 11:04:07

解決方案2
0 2014-12-12 10:59:32

使用BeautifulSoup提取部分字符串

問題描述

2 個解決方案

解決方案1 2 已采納 2014-12-12 11:04:07

解決方案2 0 2014-12-12 10:59:32

解決方案1
2 已采納 2014-12-12 11:04:07

解決方案2
0 2014-12-12 10:59:32