使用 Selenium 抓取 LinkedIn 个人资料信息

Question

I am trying to scrape profiles from LinkedIn, I get profile URLs from the below code and want to pass it to driver.get(URL), however when I scrape URLs the format of URLs is different, eg it is in [ ] brackets and I get this error我正在尝试从 LinkedIn 抓取个人资料，我从下面的代码中获取个人资料 URL，并希望将其传递给 driver.get(URL)，但是当我抓取 URL 时，URL 的格式是不同的，例如它在 [] 括号中和我收到这个错误

selenium.common.exceptions.InvalidArgumentException: Message: invalid argument: 'url' must be a string selenium.common.exceptions.InvalidArgumentException：消息：无效参数：“url”必须是字符串

Could you please suggest how to get the proper format of URLs in the list linklist = [ ] so I can pass them to driver.get(URL) .您能否建议如何在列表linklist = [ ]中获取正确格式的 URL，以便我可以将它们传递给driver.get(URL) 。 Thanks!谢谢！

options = Options()
options.add_argument("--start-maximized")
options.headless = True


url = "https://www.linkedin.com/login?fromSignIn=true&trk=guest_homepage-basic_nav-header-signin"
driver = webdriver.Chrome(path, options=options)

driver.get(url)
driver.find_element_by_id('username').send_keys('name')
driver.find_element_by_id('password').send_keys('password', Keys.ENTER)
driver.implicitly_wait(10)
driver.find_element_by_class_name('search-global-typeahead__input').send_keys('Marketing manager', Keys.ENTER)
driver.implicitly_wait(10)
driver.find_element_by_xpath('//button[text()="People"]').click()


x = 0
profile = []
linklist = []
condition = True
while condition:
    sleep(2)
    driver.execute_script("window.scrollTo(0, 1400);")
    driver.implicitly_wait(10)
    linkedin_members = driver.find_elements_by_xpath('//span[@class="entity-result__title"]')
    links = [linkedin_member.find_element_by_xpath('.//a[@class="app-aware-link"]').get_attribute('href') for linkedin_member in linkedin_members if "/in/" in linkedin_member.find_element_by_xpath('.//a[@class="app-aware-link"]').get_attribute('href')]

    x = x + 1
    linklist.append(link for link in links)
    driver.implicitly_wait(10)
    driver.find_element_by_xpath("""//button[@class='artdeco-pagination__button artdeco-pagination__button--next artdeco-button artdeco-button--muted artdeco-button--icon-right artdeco-button--1 artdeco-button--tertiary ember-view' and contains(.,'Next')]""").click()
    if x == 2:
        condition = False

profile = []

for l in tqdm(linklist):
    driver.get(l)

Answer 1

I used a for loop instad of the while loop you used, because there are no variable condition, you only want to do the loop twice.我使用for循环代替您使用的while循环，因为没有可变条件，您只想执行两次循环。

Here's how you can do it:以下是您的操作方法：

linklist = []
for i in range(2):
    time.sleep(2)
    driver.execute_script("window.scrollTo(0, 1400);")
    driver.implicitly_wait(10)
    linkedin_members = driver.find_elements_by_xpath('//span[@class="entity-result__title"]')
    
    link = driver.find_element_by_class_name('app-aware-link').get_attribute('href')
    linklist.append(link)
    driver.implicitly_wait(10)
    driver.find_element_by_xpath("""//button[@class='artdeco-pagination__button artdeco-pagination__button--next artdeco-button artdeco-button--muted artdeco-button--icon-right artdeco-button--1 artdeco-button--tertiary ember-view' and contains(.,'Next')]""").click()

for url in linklist:
    driver.get(url)

I searched the class that contains the profile url and used " .get_attribute('href') " to extract the url.我搜索了包含配置文件 url 的 class 并使用“ .get_attribute('href') ”来提取 url。

使用 Selenium 抓取 LinkedIn 个人资料信息

问题描述

1 个解决方案

解决方案1
0 2021-02-03 20:21:10

使用 Selenium 抓取 LinkedIn 个人资料信息

问题描述

1 个解决方案

解决方案1 0 2021-02-03 20:21:10

解决方案1
0 2021-02-03 20:21:10