繁体   English   中英

Python Web Scraper打印问题

[英]Python Web Scraper print issue

我已经在python中创建了一个Web抓取工具,但是在最后打印时,我想打印的代码(“ Bakerloo:” + info_from_website)已下载,就像您在代码中看到的那样,但它总是以info_from_website的形式出现,而忽略了“ Bakerloo:“字符串。 无论如何都找不到解决方案。

import urllib
import urllib.request
from bs4 import BeautifulSoup
import sys

url = 'https://tfl.gov.uk/tube-dlr-overground/status/'
page = urllib.request.urlopen(url)
soup = BeautifulSoup(page,"html.parser")

try:
   bakerlooInfo = (soup.find('li',{"class":"rainbow-list-item bakerloo "}).find_all('span')[2].text)
except:
   bakerlooInfo = (soup.find('li',{"class":"rainbow-list-item bakerloo disrupted expandable "}).find_all('span')[2].text)

bakerloo = bakerlooInfo.replace('\n','')
print("Bakerloo     : " + bakerloo)

我会改用CSS选择器 ,使元素具有disruption-summary类:

import requests
from bs4 import BeautifulSoup

url = 'https://tfl.gov.uk/tube-dlr-overground/status/'
page = requests.get(url)
soup = BeautifulSoup(page.content, "html.parser")

service = soup.select_one('li.bakerloo .disruption-summary').get_text(strip=True)
print("Bakerloo: " + service)

打印:

Bakerloo: Good service

(在此处使用requests )。


请注意,如果您只想列出所有带有干扰摘要的电台,请执行以下操作:

import requests
from bs4 import BeautifulSoup

url = 'https://tfl.gov.uk/tube-dlr-overground/status/'
page = requests.get(url)
soup = BeautifulSoup(page.content, "html.parser")

for station in soup.select("#rainbow-list-tube-dlr-overground-tflrail-tram ul li"):
    station_name = station.select_one(".service-name").get_text(strip=True)
    service_info = station.select_one(".disruption-summary").get_text(strip=True)

    print(station_name + ": " + service_info)

打印:

Bakerloo: Good service
Central: Good service
Circle: Good service
District: Good service
Hammersmith & City: Good service
Jubilee: Good service
Metropolitan: Good service
Northern: Good service
Piccadilly: Good service
Victoria: Good service
Waterloo & City: Good service
London Overground: Good service
TfL Rail: Good service
DLR: Good service
Tram: Good service

暂无
暂无

声明:本站的技术帖子网页,遵循CC BY-SA 4.0协议,如果您需要转载,请注明本站网址或者原文地址。任何问题请咨询:yoyou2525@163.com.

 
粤ICP备18138465号  © 2020-2024 STACKOOM.COM