R中的文本挖掘-從文本文件中刪除以關鍵字開頭的行

Question

我正在將文本文件讀入R，如下所示：

test<-readLines("D:/AAPL MSFT Earnings Calls/Test/Test.txt")

該文件是從PDF轉換而來的，保留了一些我想擺脫的標頭數據。 它們將以諸如“頁面”，“市值”之類的詞開頭。

如何刪除TXT文件中以這些關鍵字開頭的所有行？ 這與刪除包含該單詞的行相反。

使用以下答案之一，我修改了一些內容以閱讀

setwd("C:/Users/George/Google Drive/PhD/Strategic agility/Source Data/Peripherals Earnings Calls 2016")
text1<-readLines("test.txt")
text

library(purrr)
library(stringr)
text1 <- "foo
Page, bar
baz
Market Cap, qux"
text1 <- readLines(con = textConnection(file))
ignore_patterns <- c("^Page,", "^Market\\s+Cap,")
text1 %>% discard(~ any(str_detect(.x, ignore_patterns)))

text1

這是我得到的輸出：

> text1
[1] "foo"             "Page, bar"       "baz"             "Market Cap, qux"

foo / baz / qux字符是什么？ 謝謝

Answer 1

# once you have read and stored in a data.frame
# perform below subsetting :
x = grepl("^(Page|Market Cap)", df$id) # where df is you data.frame and 'id' is your 
                                       # column name that has those unwanted keywords
df <- df[!x,]  # does the job!

^有助於檢查開始情況。 因此，如果行以Page或（ | ） Market Cap開頭，則grepl返回TRUE

Answer 2

library(purrr)
library(stringr)
file <- "foo
Page, bar
baz
Market Cap, qux"
test <- readLines(con = textConnection(file))
ignore_patterns <- c("^Page,", "^Market\\s+Cap,")
test %>% discard(~ any(str_detect(.x, ignore_patterns)))

R中的文本挖掘-從文本文件中刪除以關鍵字開頭的行

問題描述

2 個解決方案

解決方案1
1 已采納 2017-01-10 18:27:47

解決方案2
0 2017-01-10 18:13:57

R中的文本挖掘-從文本文件中刪除以關鍵字開頭的行

問題描述

2 個解決方案

解決方案1 1 已采納 2017-01-10 18:27:47

解決方案2 0 2017-01-10 18:13:57

解決方案1
1 已采納 2017-01-10 18:27:47

解決方案2
0 2017-01-10 18:13:57