This guide shows how to extract text from a PDF document with GroupDocs.Parser for Node.js via Java.
Before you start, install the @groupdocs/groupdocs.parser package.
Extract text from a document
Create app.js next to the document you want to parse, for example sample.pdf:
'use strict';constgroupdocs=require('@groupdocs/groupdocs.parser');// Create an instance of the Parser class for the document
constparser=newgroupdocs.Parser("sample.pdf");try{// Extract the text into a reader
constreader=parser.getText();// Print the text, or a message if text extraction isn't supported for this format
console.log(reader===null?"Text extraction isn't supported":reader.readToEnd());if(reader!==null){reader.close();}}finally{// Release the document
parser.close();}process.exit(0);
Run the script:
node app.js
The script prints the text of the document:
Lorem
Lorem ipsum dolor sit amet, consectetuer adipiscing elit. Maecenas porttitor congue massa. Fusce posuere,
magna sed pulvinar ultricies, purus lectus malesuada libero, sit amet commodo magna eros quis urna.
...
Note
GroupDocs.Parser runs inside a Java virtual machine started by the java bridge, and the JVM keeps the Node.js process alive after your code finishes. Call process.exit() at the end of standalone scripts. Without a license, the extracted text is limited and contains evaluation marks; see Licensing and evaluation.
How it works
Parser opens the document. Always call close() when you are done to release the file.
getText() returns a reader with the document text, or null if text extraction isn’t supported for the document format.
Methods of GroupDocs.Parser for Node.js via Java have the same names as in GroupDocs.Parser for Java and are called synchronously.