繁体   English   中英

无法从html获取链接-jsoup

[英]Unable to get links from html - jsoup

使用以下代码,我可以从网站上获取所需的文本,但无法获取文本的关联链接。 尝试了几种方法排列和组合。 最多我得到的是下面给出的整个外部html:

<li class="list-item">
<h4><a class="bold" href="abacavir.htm">Abacavir </a>   </h4>

Abacavir is an antiviral drug that is effective against the HIV-1 virus.</li>

这是代码:

   public static void main(String[] args) throws Exception {
        Map<String,String> drugLinks = new LinkedHashMap<String,String>();
        final int OK = 200;
        //String currentURL;
        //int page = 1;
        int status = OK;
        Connection.Response response = null;
        Document doc = null;
        String[] keywords = {"a","b","c","d","e","f","g","h","i","j","k","l","m","n","o","p","q","r","s","t","u","v","w","x","y","z"};
        //String keyword = "a";
        for (String keyword : keywords){
            final String url = "https://www.medindia.net/doctors/drug_information/home.asp?alpha=" + keyword;
                response = Jsoup.connect(url)
                        .userAgent("Mozilla/5.0")
                        .execute();
                status = response.statusCode();

                    doc = response.parse();


                            Element tds = doc.select("div.related-links.top-gray.col-list.clear-fix").first();

                            Elements links = tds.select("li[class=list-item]");

                                for (Element link : links){
                                    System.out.println("generic::"+link.select("a[href]").text());
                                    System.out.println("link::"+link.attr("abs:a"));
                }

            }
        }

输出量

generic::Abacavir
link::
generic::Abacavir Sulfate and Lamivudine
link::
generic::Abacavir Sulfate, Lamivudine and Zidovudine
link::
generic::Abaloparatide
link::
generic::Abarelix
link::

如何从给定的HTML获取绝对链接?

要从元素获取链接,可以使用:

link.select("a").attr("href")

但是,这只会给您相对链接。 完整链接为:

"https://www.medindia.net/doctors/drug_information/" + link.select("a").attr("href")

暂无
暂无

声明:本站的技术帖子网页,遵循CC BY-SA 4.0协议,如果您需要转载,请注明本站网址或者原文地址。任何问题请咨询:yoyou2525@163.com.

 
粤ICP备18138465号  © 2020-2024 STACKOOM.COM