GroupDocs.Parser provides the functionality to extract formatted text from documents by the getFormattedText method:
parser.getFormattedText(options);// options is a FormattedTextOptions object
The method returns an instance of the TextReader class with the extracted text, or null if formatted text extraction isn’t supported for the document. FormattedTextOptions has the following constructor:
newgroupdocs.FormattedTextOptions(mode);// mode is a FormattedTextMode value
Reads a line of characters from the text reader and returns the data as a string.
readToEnd()
Reads all characters from the current position to the end of the text reader and returns them as one string.
close()
Releases the resources used by the reader.
Here are the steps to extract HTML formatted text from the document:
Instantiate the Parser object for the initial document;
Instantiate FormattedTextOptions with the HTML text mode;
Call the getFormattedText method and obtain the TextReader object;
Check if reader isn’t null (formatted text extraction is supported for the document);
Read the text from reader.
The following example shows how to extract a document text as HTML:
constgroupdocs=require('@groupdocs/groupdocs.parser');// Create an instance of Parser class
constparser=newgroupdocs.Parser('sample.docx');try{// Extract a formatted text into the reader
constreader=parser.getFormattedText(newgroupdocs.FormattedTextOptions(groupdocs.FormattedTextMode.Html));// If formatted text extraction isn't supported, a reader is null
if(reader==null){console.log("Formatted text extraction isn't supported");}else{try{// Print a formatted text from the document
console.log(reader.readToEnd());}finally{reader.close();}}}finally{parser.close();}process.exit(0);