用正則表達式在Python中對字符串進行分區

Question

我需要在保持空白的同時將字符串拆分為字邊界（空白）上的數組。

例如：

'this is  a\nsentence'

會成為

['this', ' ', 'is', '  ', 'a' '\n', 'sentence']

我知道str.partition和re.split ，但他們都沒有做我想要的事情，也沒有re.partition 。

我應該如何以合理的效率在Python中的空格上分區字符串？

Answer 1

試試這個：

s = "this is  a\nsentence"
re.split(r'(\W+)', s) # Notice parentheses and a plus sign.

結果將是：

['this', ' ', 'is', '  ', 'a', '\n', 'sentence']

Answer 2

re中的空格符號是'\\ s'而不是'\\ W'

相比：

import re


s = "With a sign # written @ the beginning , that's  a\nsentence,"\
    '\nno more an instruction!,\tyou know ?? "Cases" & and surprises:'\
    "that will 'lways unknown **before**, in 81% of time$"


a = re.split('(\W+)', s)
print a
print len(a)
print

b = re.split('(\s+)', s)
print b
print len(b)

產生

['With', ' ', 'a', ' ', 'sign', ' # ', 'written', ' @ ', 'the', ' ', 'beginning', ' , ', 'that', "'", 's', '  ', 'a', '\n', 'sentence', ',\n', 'no', ' ', 'more', ' ', 'an', ' ', 'instruction', '!,\t', 'you', ' ', 'know', ' ?? "', 'Cases', '" & ', 'and', ' ', 'surprises', ':', 'that', ' ', 'will', " '", 'lways', ' ', 'unknown', ' **', 'before', '**, ', 'in', ' ', '81', '% ', 'of', ' ', 'time', '$', '']
57

['With', ' ', 'a', ' ', 'sign', ' ', '#', ' ', 'written', ' ', '@', ' ', 'the', ' ', 'beginning', ' ', ',', ' ', "that's", '  ', 'a', '\n', 'sentence,', '\n', 'no', ' ', 'more', ' ', 'an', ' ', 'instruction!,', '\t', 'you', ' ', 'know', ' ', '??', ' ', '"Cases"', ' ', '&', ' ', 'and', ' ', 'surprises:that', ' ', 'will', ' ', "'lways", ' ', 'unknown', ' ', '**before**,', ' ', 'in', ' ', '81%', ' ', 'of', ' ', 'time$']
61

Answer 3

試試這個：

re.split('(\W+)','this is  a\nsentence')

用正則表達式在Python中對字符串進行分區

問題描述

3 個解決方案

解決方案1
14 已采納 2011-05-09 03:09:11

解決方案2
4 2011-05-09 06:28:49

解決方案3
3 2011-05-09 03:10:42

用正則表達式在Python中對字符串進行分區

問題描述

3 個解決方案

解決方案1 14 已采納 2011-05-09 03:09:11

解決方案2 4 2011-05-09 06:28:49

解決方案3 3 2011-05-09 03:10:42

解決方案1
14 已采納 2011-05-09 03:09:11

解決方案2
4 2011-05-09 06:28:49

解決方案3
3 2011-05-09 03:10:42