Extract hyperlinks from document page area
Leave feedback
On this page
GroupDocs.Parser provides the functionality to extract hyperlinks from a document page area by the getHyperlinks(PageAreaOptions) and getHyperlinks(int, PageAreaOptions) methods:
These methods return a Java Iterable collection of PageHyperlinkArea objects:
Member
Description
getPage()
The page that contains the hyperlink.
getRectangle()
The rectangular area on the page that contains the hyperlink.
getText()
The hyperlink text.
getUrl()
The hyperlink URL.
Here are the steps to extract hyperlinks from the document page area:
Instantiate the Parser object for the initial document;
Check if the document supports hyperlink extraction;
Instantiate PageAreaOptions with the rectangular area;
Call the getHyperlinks(PageAreaOptions) method and obtain the collection of PageHyperlinkArea objects;
Iterate through the collection and get a hyperlink text and URL.
The following example shows how to extract hyperlinks from the document page area:
constgroupdocs=require('@groupdocs/groupdocs.parser');// Create an instance of Parser class
constparser=newgroupdocs.Parser('Hyperlinks.pdf');try{// Check if the document supports hyperlink extraction
if(!parser.getFeatures().isHyperlinks()){console.log("Document doesn't support hyperlink extraction.");}else{// Create the options which are used for hyperlink extraction
constoptions=newgroupdocs.PageAreaOptions(newgroupdocs.Rectangle(newgroupdocs.Point(380,90),newgroupdocs.Size(150,50)));// Extract hyperlinks from the document page area
constit=parser.getHyperlinks(options).iterator();// Iterate over hyperlinks
while(it.hasNext()){consth=it.next();// Print the hyperlink text
console.log(h.getText());// Print the hyperlink URL
console.log(h.getUrl());console.log();}}}finally{parser.close();}process.exit(0);