HTML

The following example shows how to extract HTML formatted text:

const groupdocs = require('@groupdocs/groupdocs.parser');

// Create an instance of Parser class
const parser = new groupdocs.Parser('sample.docx');
try {
  // Extract a formatted text into the reader
  const reader = parser.getFormattedText(new groupdocs.FormattedTextOptions(groupdocs.FormattedTextMode.Html));
  // If formatted text extraction isn't supported, a reader is null
  if (reader == null) {
    console.log("Formatted text extraction isn't supported");
  } else {
    try {
      // Print a formatted text from the document
      console.log(reader.readToEnd());
    } finally {
      reader.close();
    }
  }
} finally {
  parser.close();
}
process.exit(0);
TagDescription
pParagraph is surrounded by p tag
aHyperlinks
bText with Bold font is surrounded by b tag
iText with Italic font is surrounded by i tag
h1 - h6If the heading has ‘Heading X’ style, it’s surrounded by <hX> tag
ol / ulNumbering and bullets lists
tableTables

The following Microsoft Word document is used as input document:

The following HTML document is extracted using the example above:

More resources

Free online document parser App

Along with the full-featured library we provide simple but powerful free apps.

You are welcome to extract data from PDF, DOC, DOCX, PPT, PPTX, XLS, XLSX, Emails and more with our Free Online Document Parser App.

Close
Loading

Analyzing your prompt, please hold on...

An error occurred while retrieving the results. Please refresh the page and try again.