GroupDocs.Parser allows to extract images from PDF, Emails, Ebooks, Microsoft Office: Word (DOC, DOCX), PowerPoint (PPT, PPTX), Excel (XLS, XLSX), LibreOffice formats and many others (see the full list in the supported document formats article).
GroupDocs.Parser allows to easily implement both simple and complex image extraction cases (see more in the working with images section).
In this article you can see how to extract images from any supported format without additional settings.
Extract images from documents
To extract images from documents simply call the getImages() method:
parser.getImages();// returns a Java Iterable of PageImageArea objects or null
This method returns a collection of PageImageArea objects:
Member
Description
getPage()
The page that contains the image area.
getRectangle()
The rectangular area on the page that contains the image area.
getFileType()
The format of the image.
getRotation()
The rotation angle of the image.
getImageStream()
Returns the image stream.
getImageStream(options)
Returns the image stream in a different format.
save(filePath)
Saves the image to the file.
save(filePath, options)
Saves the image to the file in a different format.
Here are the steps to extract images from the whole document:
Instantiate the Parser object for the initial document;
Call the getImages() method and obtain the collection of image objects;
Check if collection isn’t null (image extraction is supported for the document);
Iterate through the collection and get sizes, image types and image contents.
The collection is a Java Iterable, so iterate over it with iterator(), hasNext() and next().
The following example shows how to extract all images from the whole document:
constgroupdocs=require('@groupdocs/groupdocs.parser');// Create an instance of Parser class
constparser=newgroupdocs.Parser('images.pdf');try{// Extract images
constimages=parser.getImages();// Check if images extraction is supported
if(images===null){console.log("Images extraction isn't supported");}else{// Iterate over images
constit=images.iterator();while(it.hasNext()){constimage=it.next();// Print a page index, rectangle and image type
console.log(`Page: ${image.getPage().getIndex()}, R: ${image.getRectangle().toString()}, Type: ${image.getFileType().toString()}`);}}}finally{parser.close();}process.exit(0);
More resources
Advanced usage topics
To learn more about document data extraction features and get familiar how to extract text, images, forms and more, please refer to the advanced usage section.
Free online document parser App
Along with the full-featured library we provide simple, but powerful free Apps.
You are welcome to extract data from PDF, DOC, DOCX, PPT, PPTX, XLS, XLSX, Emails and more with our free online Free Online Document Parser App.
Was this page helpful?
Any additional feedback you'd like to share with us?
Please tell us how we can improve this page.
Thank you for your feedback!
We value your opinion. Your feedback will help us improve our documentation.
On this page
Analyzing your prompt, please hold on...
An error occurred while retrieving the results. Please refresh the page and try again.