GroupDocs.Parser provides the Document Parser feature that allows you to extract data from documents of various formats including PDF, Microsoft Word, Excel, LibreOffice formats etc. (see the full supported list).
With the Document Parsing feature you can easily solve business automation tasks with the data extracted from your documents.
Using this feature is straightforward. Simply define a template programmatically and apply it.
Parse data from documents
GroupDocs.Parser provides the functionality to extract data from documents by the parseByTemplate(template) method:
parser.parseByTemplate(template);// returns DocumentData or null
This method parses data from the document by a user-generated template.
Here are the steps to parse data from the document by a user-generated template:
Instantiate the Parser object for the initial document;
Instantiate the Template object with the user-generated template;
Call the parseByTemplate(template) method and obtain the DocumentData object;
Check if data isn’t null (parse by template is supported for the document);
Iterate over field data to obtain the extracted data.
The Template constructor accepts a Java collection of template items, so the example creates a java.util.ArrayList and adds the items to it. Use java.instanceOf(object, className) from the java package to check the Java type of a field value.
The following example shows how to parse data from the document by a user-generated template:
constjava=require('java');constgroupdocs=require('@groupdocs/groupdocs.parser');functionrect(x,y,width,height){returnnewgroupdocs.Rectangle(newgroupdocs.Point(x,y),newgroupdocs.Size(width,height));}functionfixedField(x,y,width,height,name){returnnewgroupdocs.TemplateField(newgroupdocs.TemplateFixedPosition(rect(x,y,width,height)),name);}functionlinkedField(linkedFieldName,name){// The value is located to the right of the linked field
constedges=newgroupdocs.TemplateLinkedPositionEdges(false,false,true,false);returnnewgroupdocs.TemplateField(newgroupdocs.TemplateLinkedPosition(linkedFieldName,newgroupdocs.Size(200,15),edges),name);}functiongetTemplate(){// Create detector parameters for "Details" table
constdetailsTableParameters=newgroupdocs.TemplateTableParameters(rect(35,320,530,55),null);// Create detector parameters for "Summary" table
constsummaryTableParameters=newgroupdocs.TemplateTableParameters(rect(330,385,220,65),null);// Create a collection of template items
consttemplateItems=[fixedField(35,135,100,10,'FromCompany'),fixedField(35,150,100,35,'FromAddress'),fixedField(35,190,150,2,'FromEmail'),fixedField(35,250,100,2,'ToCompany'),fixedField(35,260,100,15,'ToAddress'),fixedField(35,290,150,2,'ToEmail'),newgroupdocs.TemplateField(newgroupdocs.TemplateRegexPosition('Invoice Number'),'InvoiceNumber'),linkedField('InvoiceNumber','InvoiceNumberValue'),newgroupdocs.TemplateField(newgroupdocs.TemplateRegexPosition('Order Number'),'InvoiceOrder'),linkedField('InvoiceOrder','InvoiceOrderValue'),newgroupdocs.TemplateField(newgroupdocs.TemplateRegexPosition('Invoice Date'),'InvoiceDate'),linkedField('InvoiceDate','InvoiceDateValue'),newgroupdocs.TemplateField(newgroupdocs.TemplateRegexPosition('Due Date'),'DueDate'),linkedField('DueDate','DueDateValue'),newgroupdocs.TemplateField(newgroupdocs.TemplateRegexPosition('Total Due'),'TotalDue'),linkedField('TotalDue','TotalDueValue'),newgroupdocs.TemplateTable(detailsTableParameters,'details',null),newgroupdocs.TemplateTable(summaryTableParameters,'summary',null),];// Create a document template
constitems=new(java.import('java.util.ArrayList'))();templateItems.forEach((item)=>items.add(item));returnnewgroupdocs.Template(items);}// Create an instance of Parser class
constparser=newgroupdocs.Parser('invoice.pdf');try{// Parse the document by the template
constdata=parser.parseByTemplate(getTemplate());// Check if parsing by template is supported
if(data===null){console.log("Parse Document by Template isn't supported.");}else{// Print extracted fields
for(leti=0;i<data.getCount();i++){constfield=data.get(i);constarea=field.getPageArea();// Check if the field value is a text area
constisTextArea=area!==null&&java.instanceOf(area,'com.groupdocs.parser.data.PageTextArea');console.log(`${field.getName()}: ${isTextArea?area.getText():'Not a template field'}`);}}}finally{parser.close();}process.exit(0);
The beginning of the output looks like this:
FROMCOMPANY: DEMO - Sliced Invoices
FROMADDRESS: Suite 5A-1204
123 Somewhere Street
Your City AZ 12345
FROMEMAIL: admin@slicedinvoices.com
...
INVOICENUMBERVALUE: INV-3337
...
TOTALDUEVALUE: $93.50
DETAILS: Not a template field
SUMMARY: Not a template field
The details and summary fields contain tables (PageTableArea objects) rather than text areas.
More resources
Advanced usage topics
To learn more about template building and working with extracted data please refer to the following guides: