正则表达式以一致的顺序提取字符串的不同部分

Question

I have a list of strings 我有一个字符串列表

my_strings = [
    "2002-03-04 with Matt",
    "Important: 2016-01-23 with Mary",
    "with Tom on 2015-06-30",
]

I want to extract: 我想提取：

date (always in yyyy-mm-dd format) 日期（始终采用yyyy-mm-dd格式）
person (always in with person) but I don't want to keep "with" 人（总是与人在一起）但我不想保持“与”

I could do: 我可以做：

import re
pattern = r'.*(\d{4}-\d{2}-\d{2}).*with \b([^\b]+)\b.*'
matched = [re.match(pattern, x).groups() for x in my_strings]

but it fails because pattern doesn't match "with Tom on 2015-06-30" . 但它失败了，因为模式与"with Tom on 2015-06-30"不匹配。

Question s 问小号

How do I specify the regex pattern to be indifferent to the order in which date or person appear in the string? 如何指定正则表达式模式对日期或人物出现在字符串中的顺序无动于衷？

and 和

How do I ensure that the groups() method returns them in the same order every time? 如何确保groups()方法每次都以相同的顺序返回它们？

I expect the output to look like this? 我希望输出看起来像这样？

[('2002-03-04', 'Matt'), ('2016-01-23', 'Mary'), ('2015-06-30', 'Tom')]

Answer 1

What about doing it with 2 separate regex? 用2个独立的正则表达式做什么呢？

my_strings = [
    "2002-03-04 with Matt",
    "Important: 2016-01-23 with Mary",
    "with Tom on 2015-06-30",
]
import re

pattern = r'.*(\d{4}-\d{2}-\d{2})'
dates = [re.match(pattern, x).groups()[0] for x in my_strings]

pattern = r'.*with (\w+).*'
persons = [re.match(pattern, x).groups()[0] for x in my_strings]

output = zip(dates, persons)
print output
## [('2002-03-04', 'Matt'), ('2016-01-23', 'Mary'), ('2015-06-30', 'Tom')]

Answer 2

This should work: 这应该工作：

my_strings = [
    "2002-03-04 with Matt",
    "Important: 2016-01-23 with Mary",
    "with Tom on 2015-06-30",
]

import re

alternates = r"(?:\b(\d{4}-\d\d-\d\d)\b|with (\w+)|.)*"

for tc in my_strings:
    print(tc)
    m = re.match(alternates, tc)
    if m:
        print("\t", m.group(1))
        print("\t", m.group(2))

Output is: 输出是：

$ python test.py
2002-03-04 with Matt
     2002-03-04
     Matt
Important: 2016-01-23 with Mary
     2016-01-23
     Mary
with Tom on 2015-06-30
     2015-06-30
     Tom

However, something like this is not totally intuitive. 但是，这样的事情并不完全直观。 I encourage you to try using named groups if at all possible. 我鼓励你尽可能尝试使用命名组。

Answer 3

Just for education reasons, a non-regex approach could involve using dateutil parser in a "fuzzy" mode to extract the dates and the nltk toolkit with the named entity recognition to extract names. 出于教育原因，非正则表达式方法可能涉及在“模糊”模式下使用dateutil解析器来提取日期，并使用命名实体识别来提取nltk工具包以提取名称。 Complete code: 完整代码：

import nltk
from nltk import pos_tag, ne_chunk
from nltk.tokenize import SpaceTokenizer
from dateutil.parser import parse


def extract_names(text):
    tokenizer = SpaceTokenizer()
    toks = tokenizer.tokenize(text)
    pos = pos_tag(toks)
    chunked_nes = ne_chunk(pos)

    return [' '.join(map(lambda x: x[0], ne.leaves())) for ne in chunked_nes if isinstance(ne, nltk.tree.Tree)]

my_strings = [
    "2002-03-04 with Matt",
    "Important: 2016-01-23 with Mary",
    "with Tom on 2015-06-30"
]

for s in my_strings:
    print(parse(s, fuzzy=True))
    print(extract_names(s))

Prints: 打印：

2002-03-04 00:00:00
['Matt']
2016-01-23 00:00:00
['Mary']
2015-06-30 00:00:00
['Tom']

That's probably an over-complication though. 但这可能是一个过于复杂的问题。

Answer 4

If you use Python's new regex module, you can use conditionals to get 如果您使用Python的新正则表达式模块，则可以使用条件来获取
a guaranteed match on 2 items. 保证匹配2件物品。

I'd think this is more like a standard to do out-of-order matching. 我认为这更像是一个无序匹配的标准。

(?:.*?(?:(?(1)(?!))\\b(\\d{4}-\\d\\d-\\d\\d)\\b|(?(2)(?!))with[ ](\\w+))){2}

Expanded 扩展

 (?:
      .*? 
      (?:
           (?(1)(?!))
           \b 
           ( \d{4} - \d\d - \d\d )       # (1)
           \b 
        |  (?(2)(?!))
           with [ ] 
           ( \w+ )                       # (2)
      )
 ){2}

正则表达式以一致的顺序提取字符串的不同部分

问题描述

Question s 问小号

4 个解决方案

解决方案1
4 2016-05-09 19:47:50

解决方案2
2 2016-05-09 19:43:59

解决方案3
2 2016-05-09 19:51:09

解决方案4
2 已采纳 2016-05-09 20:07:28

正则表达式以一致的顺序提取字符串的不同部分

问题描述

Question s 问小号

4 个解决方案

解决方案1 4 2016-05-09 19:47:50

解决方案2 2 2016-05-09 19:43:59

解决方案3 2 2016-05-09 19:51:09

解决方案4 2 已采纳 2016-05-09 20:07:28

解决方案1
4 2016-05-09 19:47:50

解决方案2
2 2016-05-09 19:43:59

解决方案3
2 2016-05-09 19:51:09

解决方案4
2 已采纳 2016-05-09 20:07:28