Data extracted by the parseByTemplate method are stored in the instance of the DocumentData class:
Member
Description
getCount()
The total number of the data fields.
get(int)
The data field.
getFieldsByName(String)
Returns the Java list of data fields where the name is equal to fieldName.
iterator()
Returns the Java iterator over the data fields.
The FieldData class has the following members:
Member
Description
getName()
The field name.
getPageIndex()
The page index.
getPageArea()
The value of the field.
getLinkedField()
The linked field.
Field data are stored in the getPageArea() property. Depending on the type of the value it can contain an instance of the PageTextArea, PageTableArea or PageBarcodeArea classes. JavaScript instanceof doesn’t work with Java objects; check the type with java.instanceOf:
constjava=require('java');// Get the field data
constfield=data.get(i);// Check if the field data contains a text
if(java.instanceOf(field.getPageArea(),'com.groupdocs.parser.data.PageTextArea')){// Print the field value
console.log(field.getPageArea().getText());}
The PageTextArea class represents a text block on the page. This class has the following members:
Member
Description
getRectangle()
The rectangular area that bounds the text area.
getPage()
The page information (page index and page size).
getText()
The value of the text area.
getBaseLine()
The base line of the text area.
getTextStyle()
The style of the text block (like font name, font size etc.)
getAreas()
The collection of child text areas.
The text area can be single or composite. In the first case it contains a text which is bounded by a rectangular area. In the second case it contains other text areas; text and table properties are calculated by child text areas.
The PageTableArea class represents a table. This class has the following members:
Member
Description
getRectangle()
The rectangular area that bounds the table.
getPage()
The page information (page index and page size).
getRowCount()
The total number of the table rows.
getColumnCount()
The total number of the table columns.
getCell(int, int)
The table cell by row and column indexes.
getRowHeight(int)
Returns the row height.
getColumnWidth(int)
Returns the column width.
There are two ways to work with fields data.
Iterate through fields
The following example shows how to iterate via extracted field data:
constjava=require('java');constgroupdocs=require('@groupdocs/groupdocs.parser');constArrayList=java.import('java.util.ArrayList');// Define a "price" field
constpriceField=newgroupdocs.TemplateField(newgroupdocs.TemplateRegexPosition('\\$\\d+(.\\d+)?'),'Price');// Define a "email" field
constemailField=newgroupdocs.TemplateField(newgroupdocs.TemplateRegexPosition('[a-z]+\\@[a-z]+.[a-z]+'),'Email');// Create a template
constitems=newArrayList();items.add(priceField);items.add(emailField);consttemplate=newgroupdocs.Template(items);// Create an instance of Parser class
constparser=newgroupdocs.Parser('invoice.pdf');try{// Parse the document by the template
constdata=parser.parseByTemplate(template);// Print all extracted data
for(leti=0;i<data.getCount();i++){constfield=data.get(i);// As we have defined only text fields in the template,
// the page area is expected to be PageTextArea
constarea=field.getPageArea();constvalue=java.instanceOf(area,'com.groupdocs.parser.data.PageTextArea')?area.getText():'Not a template field';// Print field name and value
console.log(field.getName()+': '+value);}}finally{parser.close();}process.exit(0);
The following example shows how to get fields by the name:
constjava=require('java');constgroupdocs=require('@groupdocs/groupdocs.parser');constArrayList=java.import('java.util.ArrayList');// Define "price" and "email" fields
constitems=newArrayList();items.add(newgroupdocs.TemplateField(newgroupdocs.TemplateRegexPosition('\\$\\d+(.\\d+)?'),'Price'));items.add(newgroupdocs.TemplateField(newgroupdocs.TemplateRegexPosition('[a-z]+\\@[a-z]+.[a-z]+'),'Email'));// Create a template
consttemplate=newgroupdocs.Template(items);// Print values of the fields with the name
functionprintFields(data,name){constit=data.getFieldsByName(name).iterator();while(it.hasNext()){constarea=it.next().getPageArea();console.log(java.instanceOf(area,'com.groupdocs.parser.data.PageTextArea')?area.getText():'Not a template field');}}// Create an instance of Parser class
constparser=newgroupdocs.Parser('invoice.pdf');try{// Parse the document by the template
constdata=parser.parseByTemplate(template);// Print prices
console.log('Prices:');printFields(data,'Price');// Print emails
console.log('Emails:');printFields(data,'Email');}finally{parser.close();}process.exit(0);
This functionality allows to iterate all data fields and select the most suitable of them. For example, if more than one text value meets the condition of the regular expression, a user can iterate over them and select the most suitable one.
Working with tables
The following example shows how to work with extracted tables:
constjava=require('java');constgroupdocs=require('@groupdocs/groupdocs.parser');constArrayList=java.import('java.util.ArrayList');// Create a table template with the parameters
consttable=newgroupdocs.TemplateTable(newgroupdocs.TemplateTableParameters(newgroupdocs.Rectangle(newgroupdocs.Point(35,320),newgroupdocs.Size(530,55)),null),'Details',null);// Create a template
constitems=newArrayList();items.add(table);consttemplate=newgroupdocs.Template(items);// Create an instance of Parser class
constparser=newgroupdocs.Parser('invoice.pdf');try{// Parse the document by the template
constdata=parser.parseByTemplate(template);// Print all extracted data
for(leti=0;i<data.getCount();i++){console.log(data.get(i).getName()+':');// Check if the field is a table
constarea=data.get(i).getPageArea();if(!java.instanceOf(area,'com.groupdocs.parser.data.PageTableArea')){continue;}// Iterate via table rows
for(letrow=0;row<area.getRowCount();row++){constcells=[];// Iterate via table columns
for(letcolumn=0;column<area.getColumnCount();column++){// Get the cell value
constcellValue=area.getCell(row,column).getPageArea();cells.push(java.instanceOf(cellValue,'com.groupdocs.parser.data.PageTextArea')?cellValue.getText():'');}// Print the row; cells are separated by tabs
console.log(cells.join('\t'));}}}finally{parser.close();}process.exit(0);
More resources
Free online document parser App
Along with the full-featured library we provide simple but powerful free Apps.
You are welcome to parse documents and extract data from PDF, DOC, DOCX, PPT, PPTX, XLS, XLSX, Emails and more with our Free Online Document Parser App.
Was this page helpful?
Any additional feedback you'd like to share with us?
Please tell us how we can improve this page.
Thank you for your feedback!
We value your opinion. Your feedback will help us improve our documentation.
On this page
Analyzing your prompt, please hold on...
An error occurred while retrieving the results. Please refresh the page and try again.