Web 刮 BeautifulSoup find_all 在 APEC 上不起作用

Question

我想查找所有 div class 容器结果，但它不起作用我得到一个空列表

我的代码：`

from bs4 import BeautifulSoup, NavigableString, Tag
import requests
import urllib.request
url = "https://www.apec.fr/candidat/recherche-emploi.html/emploi?page="
for page in range(0,10,1):
    r = requests.get(url + str(page))
    soup = BeautifulSoup(r.content,"html.parser")
    ancher = soup.find_all('div', attrs={'class': 'container-result'})
print(ancher)

`

Answer 1

由于 web 页面是由 javascript 渲染的，因此请求 / BeautifulSoup 将无法检索所需的 DOM 元素，因为它们是在渲染页面的一段时间后添加的。 为此，您可以尝试使用 selenium，这是一个示例：

from selenium import webdriver 
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait

# delay for selenium web driver wait
DELAY = 30

# create selenium driver
chrome_options = webdriver.ChromeOptions()
#chrome_options.add_argument('--headless')
#chrome_options.add_argument('--no-sandbox')
driver = webdriver.Chrome('<<PATH TO chromedriver>>', options=chrome_options)

# iterate over pages
for page in range(0, 10, 1):
    
    # open web page
    driver.get(f'https://www.apec.fr/candidat/recherche-emploi.html/emploi?page={page}')
    
    # wait for element with class 'container-result' to be added
    container_result = WebDriverWait(driver, DELAY).until(EC.presence_of_element_located((By.CLASS_NAME, "container-result")))
    # scroll to container-result
    driver.execute_script("arguments[0].scrollIntoView();", container_result)
    # get source HTML of the container-result element
    source = container_result.get_attribute('innerHTML')
    # print source
    print(source)
    # here you can continue work with the source variable either using selenium API or using BeautifulSoup API:
    # soup = BeautifulSoup(source, "html.parser")

# quit webdriver    
driver.quit()

Web 刮 BeautifulSoup find_all 在 APEC 上不起作用

问题描述

1 个解决方案

解决方案1
0 已采纳 2021-01-09 11:21:47

Web 刮 BeautifulSoup find_all 在 APEC 上不起作用

问题描述

1 个解决方案

解决方案1 0 已采纳 2021-01-09 11:21:47

解决方案1
0 已采纳 2021-01-09 11:21:47