如何使用Perl提取HTML文件的特定部分

Question

我是Perl的新手，我正在嘗試讀取HTML文件的<div class="one">之間的特定內容。

HTML檔案：

<div class="one">

    <div id="two">Donec eu libero sit amet quam egestas semper. Aenean ultricies mi vitae est. Mauris placerat eleifend leo.
    </div>

    <pre>Pellentesque habitant morbi tristique senectus et netus et malesuada fames ac turpis egestas.
    </pre>

</div>

Perl代碼：

my $file = "content.html";

if (-e $file) {
    open(IN, $file);
    while (<IN>) {
        chomp($line = $_);

        #print "$line\n";
    }
}

@contents = <IN>;

#check to if content in html file is in the right location,
#if content is in correct location (div class="one")
#print content in div two and three if exist

for (my $i = 0 ; $i <= $#contents ; $i++) {
    if (!$contents[$i] =~ m/^\s*<div/ && $contents[$i] =~ m/class\s*=\s*"one"/) {
        print "content in wrong location";
    }
    else {
        if ($contents[$i] =~ m/^\s*<div/) {
            print "$_";
        }
        else ($contents[$i] =~ m/^\s*<pre/) {
            print "$_";
        }
    }
}

Answer 1

使用HTML :: TreeBuilder取得了一些成功，它擅長處理損壞的HTML。

如何使用Perl提取HTML文件的特定部分

問題描述

1 個解決方案

解決方案1
1 2013-04-22 17:57:07

如何使用Perl提取HTML文件的特定部分

問題描述

1 個解決方案

解決方案1 1 2013-04-22 17:57:07

解決方案1
1 2013-04-22 17:57:07