繁体   English   中英

当使用urlparser在日志文件中循环访问某些行时缺少特定参数时,如何在列表中附加null

[英]How can I append null in list when particular parameter is absent for some lines while iterating through log file using urlparser

我希望通过url来解析我的文件,但是某些url缺少参数,并且当我遍历日志行时遇到了缺少参数的错误。 我需要将空白或空值添加到解析列表中,以便可以将其转换为数据框

我的数据文件:日志文件

"GET /pixel.gife=heartbeat&creative_id=33548&in_view_time=290"
"GET/pixel.gife=heartbeat&creative_id=33548&in_view_time=23988"
"GET /pixel.gif?e=heartbeat&creative_id=33548&in_view_time=19183"
"GET /pixel.gif?e=ad_load&creative_id=33548"

我希望输出为:

   E |  Creative ID | IN VIEW TIME

   heartbeat   33548    290

   heartbeat 33548 23988

   ad_load 33548 null

我的代码:

parselist = []
for eachline in log.readlines():
    ip_regex = re.findall(r'(\d{18})', eachline)
    date = re.findall(r'([0-9]{4}\-[0-9]{2}\-[0-9]{2})',eachline)
    url = eachline
    parsed = urlparse.urlparse(url)
    parselist.append(ip_regex)
    parselist.append(date)
    parselist.append(urlparse.parse_qs(parsed.query)['e'])
    parselist.append(urlparse.parse_qs(parsed.query)['account_id'])
    parselist.append(urlparse.parse_qs(parsed.query)['impression_id'])
    parselist.append(urlparse.parse_qs(parsed.query)['campaign_id'])
    parselist.append(urlparse.parse_qs(parsed.query)['creative_id'])
    parselist.append(urlparse.parse_qs(parsed.query)['in_view_time'])

我收到错误,因为第三行缺少in_view_time参数:

---------------------------------------------------------------------------
KeyError                                  Traceback (most recent call last)
<ipython-input-6-405c1bfb329e> in <module>()
     12     parselist.append(urlparse.parse_qs(parsed.query)['campaign_id'])
     13     parselist.append(urlparse.parse_qs(parsed.query)['creative_id'])
---> 14     parselist.append(urlparse.parse_qs(parsed.query)['in_view_time'])

KeyError: 'in_view_time'

您可以使用tryexcept

parselist = []
for eachline in log.readlines():
    ip_regex = re.findall(r'(\d{18})', eachline)
    date = re.findall(r'([0-9]{4}\-[0-9]{2}\-[0-9]{2})',eachline)
    url = eachline
    parsed = urlparse.urlparse(url)
    parselist.append(ip_regex)
    parselist.append(date)
    try:
        parselist.append(urlparse.parse_qs(parsed.query)['e'])
    except:
        parselist.append('Null')
    try:
        parselist.append(urlparse.parse_qs(parsed.query)['account_id'])
    except:
        parselist.append('Null')
    try:
        parselist.append(urlparse.parse_qs(parsed.query)['impression_id'])
    except:
        parselist.append('Null')
    try:
        parselist.append(urlparse.parse_qs(parsed.query)['campaign_id'])
    except:
        parselist.append('Null')
    try:
        parselist.append(urlparse.parse_qs(parsed.query)['creative_id'])
    except:
        parselist.append('Null')
    try:
        parselist.append(urlparse.parse_qs(parsed.query)['in_view_time'])
    except:
        parselist.append('Null')

或者,以更紧凑的方式:

parselist = []
for eachline in log.readlines():
    ip_regex = re.findall(r'(\d{18})', eachline)
    date = re.findall(r'([0-9]{4}\-[0-9]{2}\-[0-9]{2})',eachline)
    url = eachline
    parsed = urlparse.urlparse(url)
    parselist.append(ip_regex)
    parselist.append(date)

    for key in ['e','account_id','impression_id','campaign_id','creative_id','in_view_time']:
        try:
            parselist.append(urlparse.parse_qs(parsed.query)[key])
        except:
            parselist.append('Null')

作为建议,您可以附加None来代替'Null'

  1. 为什么要创建列表(丢失密钥并仅存储值)?
  2. 如果您只对这些值感兴趣,则可以简单地编写以下内容:
for v in urlparse.parse_qs(parsed.query).values():
    parselist.append(v)

暂无
暂无

声明:本站的技术帖子网页,遵循CC BY-SA 4.0协议,如果您需要转载,请注明本站网址或者原文地址。任何问题请咨询:yoyou2525@163.com.

 
粤ICP备18138465号  © 2020-2024 STACKOOM.COM