GroupDocs.Parser provides the functionality to extract images from a document page by the getImages(int) method:
parser.getImages(pageIndex)// returns Iterable<PageImageArea> or null
This method returns a Java Iterable collection of PageImageArea objects:
Member
Description
getPage()
The page that contains the image.
getRectangle()
The rectangular area on the page that contains the image.
getFileType()
The format of the image.
getRotation()
The rotation angle of the image.
getImageStream()
Returns the image stream (java.io.InputStream).
getImageStream(ImageOptions)
Returns the image stream in a different format.
save(String)
Saves the image to the file.
save(String, ImageOptions)
Saves the image to the file in a different format.
The ImageOptions class is used to define the image format into which the image is converted. The following image formats (ImageFormat) are supported:
Bmp
Gif
Jpeg
Png
WebP
Here are the steps to extract images from the document page:
Instantiate the Parser object for the initial document;
Call parser.getFeatures().isImages() to check if images extraction is supported for the document;
Call the getImages(int) method with the page index and obtain the collection of PageImageArea objects;
Iterate through the collection and get sizes, image types and image contents.
The following example shows how to extract images from each document page:
constgroupdocs=require('@groupdocs/groupdocs.parser');// Create an instance of Parser class
constparser=newgroupdocs.Parser('images.pdf');try{// Check if the document supports images extraction
if(!parser.getFeatures().isImages()){console.log("Document doesn't support images extraction.");}else{// Get the document info
constdocumentInfo=parser.getDocumentInfo();constpageCount=documentInfo.getPageCount();// Iterate over pages
for(letpageIndex=0;pageIndex<pageCount;pageIndex++){// Print a page number
console.log(`Page ${pageIndex+1}/${pageCount}`);// Iterate over images
// We ignore null-checking as we have checked images extraction feature support earlier
constit=parser.getImages(pageIndex).iterator();while(it.hasNext()){constimage=it.next();// Print a rectangle and image type
console.log(`R: ${image.getRectangle()}, Type: ${image.getFileType()}`);}}}}finally{parser.close();}process.exit(0);