[英]Python & lxml / xpath: Parsing XML
我需要从此链接获取FLVPath的值: http : //www.testpage.com/v2/videoConfigXmlCode.php? ppg = video_29746_no_0_extsite
from lxml import html
sub_r = requests.get("http://www.testpage.co/v2/videoConfigXmlCode.php?pg=video_%s_no_0_extsite" % list[6])
sub_root = lxml.html.fromstring(sub_r.content)
for sub_data in sub_root.xpath('//PLAYER_SETTINGS[@Name="FLVPath"]/@Value'):
print sub_data.text
但没有数据返回
您正在使用lxml.html
来解析文档,这会导致lxml小写所有元素和属性名称(因为这在html中无关紧要),这意味着您必须使用:
sub_root.xpath('//player_settings[@name="FLVPath"]/@value')
或者当您正在解析xml文件时,您可以使用lxml.etree
。
你可以试试
print sub_data.attrib['Value']
url = "http://www.testpage.com/v2/videoConfigXmlCode.php?pg=video_29746_no_0_extsite"
response = requests.get(url)
# Use `lxml.etree` rathern than `lxml.html`,
# and unicode `response.text` instead of `response.content`
doc = lxml.etree.fromstring(response.text)
for path in doc.xpath('//PLAYER_SETTINGS[@Name="FLVPath"]/@Value'):
print path
声明:本站的技术帖子网页,遵循CC BY-SA 4.0协议,如果您需要转载,请注明本站网址或者原文地址。任何问题请咨询:yoyou2525@163.com.