在Python中使用BS4解析HTML

Question

我正在嘗試使用以下HTML解析網站。

我正在使用Python和BeautifulSoup。

如何從中提取德州游騎兵的文字？

我不在上課，所以遇到了麻煩？ 謝謝，

馬特

<div class="team">
            <span class="team-logo mlb tex"></span>Texas Rangers
                            <br />
                <a class="fancy" href="/split_stats/index/Baseball/Pitcher/107">BvP</a>
                &middot;


                                <a class="fancy" href="/split_stats/index/Baseball/Righty/107">vs. R/a>
                &middot;

                <a class="fancy" href="/split_stats/index/Baseball/Away/107">Away</a>
                &middot;

                                <a class="fancy" href="/split_stats/index/Baseball/Night/107">Night</a>

                    </div>

Answer 1

可能不是最好的解決方案，但這可行。

>>> soup = BeautifulSoup(htmlCode)
>>> soup.div.contents[2].strip()
u'Texas Rangers'

Answer 2

我將使用在ipython中運行的以下代碼：

In [28]: htmldoc = """<div class="team">
   ....: <span class="team-logo mlb tex"></span>Texas Rangers
   ....: <br />
   ....: <a class="fancy" href="/split_stats/index/Baseball/Pitcher/107">BvP</a>
   ....: &middot;
   ....: <a class="fancy" href="/split_stats/index/Baseball/Righty/107">vs. R/a&gt;
   ....: &middot;
   ....: </a><a class="fancy" href="/split_stats/index/Baseball/Away/107">Away</a>
   ....: &middot;
<   ....: <a class="fancy" href="/split_stats/index/Baseball/Night/107">Night</a>
   ....: </div>
   ....: """

In [30]: soup = BeautifulSoup(htmldoc)

In [31]: import re

In [32]: soup(text=re.compile('Texas Rangers'))
Out[32]: [u'Texas Rangers\n']

在Python中使用BS4解析HTML

問題描述

2 個解決方案

解決方案1
2 2014-07-23 03:05:54

解決方案2
0 2014-07-23 03:15:29

在Python中使用BS4解析HTML

問題描述

2 個解決方案

解決方案1 2 2014-07-23 03:05:54

解決方案2 0 2014-07-23 03:15:29

解決方案1
2 2014-07-23 03:05:54

解決方案2
0 2014-07-23 03:15:29