使用Selenium Webdriver（Python）從網站提取圖像

Question

我需要抓取數千個子站點並提取信息。

現在，不幸的是，所討論的信息不是常規的HTML文本，而是動態呈現文本的圖像。

如何提取這些圖像以進一步處理它們？ 我在Python上使用Selenium Webdriver。

Answer 1

mechanize加上BeautifulSoup幾乎是無法完成的。 圖像的進一步處理可以使用pytesser完成，但是我在那里沒有經驗。 有經驗的人提供有關Python OCR知識的建議會很有趣。

導入機械化，BeautifulSoup

browser = mechanize.Browser()
html = browser.open("http://www.dreamstime.com/free-photos")
soup = BeautifulSoup.BeautifulSoup(html)
for ii, image in enumerate(soup.findAll('img')):
    _src = image['src']
    if str(_src).startswith('http://') and str(_src).endswith('.jpg'):
        print 'Storing this image:', _src
        data = browser.open(_src).read()
        fl = 'image' + str(ii) + '.jpg'
        with open(fl, 'wb') as f:
            f.write(data)
        f.closed

使用Selenium Webdriver（Python）從網站提取圖像

問題描述

1 個解決方案

解決方案1
0 2013-09-02 10:03:42

使用Selenium Webdriver（Python）從網站提取圖像

問題描述

1 個解決方案

解決方案1 0 2013-09-02 10:03:42

解決方案1
0 2013-09-02 10:03:42