R 更快的替代 str_contains

Question

I've got a table and a character vector:我有一个表格和一个字符向量：

ACT_Suburbs_Names <- data.table(DivisionNm = c('ACTON', 'AINSLIE', 'AMAROO', 'ARANDA', 'BANKS'))

temp1 <- c('U1 336 DOORING ST ACTON', '65/78 ACHELON DR MOORONG')

The script below checks that the addresses in temp1 are in the list of addresses from ACT_Suburbs_Names.下面的脚本检查 temp1 中的地址是否在 ACT_Suburbs_Names 的地址列表中。

future_sapply(temp1, function(x) str_contains(ACT_Suburbs_Names, x, ignore.case = TRUE))

It takes an enormous amount of time to process all the data that I've got.处理我拥有的所有数据需要大量时间。 Anything faster?有更快的吗？ even in python would be ok.即使在 python 中也可以。

Answer 1

Create a pattern and use regular expressions:创建一个模式并使用正则表达式：

> library(tidyverse)
> 
> ACT_Suburbs_Names <-  c('ACTON', 'AINSLIE', 'AMAROO', 'ARANDA', 'BANKS')
> 
> pat <- paste(ACT_Suburbs_Names, collapse = '|')
> 
> temp1 <- c('U1 336 DOORING ST ACTON', 
+  '65/78 ACHELON DR MOORONG',
+  'asdf ARANDA asdfsadf',
+           ' asdfasfd   BANK asdf',
+           ' sdafasdf  BANKS  asdf',
+           ' just junk') 
> 
> # find out which entries match - return the index of a match
> 
> grep(pat, temp1)
[1] 1 3 5
>

R 更快的替代 str_contains

问题描述

1 个解决方案

解决方案1
2 已采纳 2020-08-18 02:39:16

R 更快的替代 str_contains

问题描述

1 个解决方案

解决方案1 2 已采纳 2020-08-18 02:39:16

解决方案1
2 已采纳 2020-08-18 02:39:16