PDF Extract data function

The PDF Extract data function in RPA.

Help CentreBuilding Processes

This article describes the RPA PDF functionality to extract data.

For a full list of available functions, please read PDF functions.


Description

The PDF Extract data function allows you to extract data.


Setting it up

To use the PDF Extract data function, follow these instructions:

{{popups:rpas/add:action=pdf-extract}}

Once on a process map:

  1. Navigate to the process where you would like to add the PDF Extract data function
  2. Click on the green button’s drop-down (right side of the green button labelled “Process”)
  3. On the menu that appears, click on the “+ Activity” item
  4. From here you can select the PDF Extract data option
  5. Click "Add" once you are done

Input parameters

ParameterRequiredTypeDefaultDescription
FileYesText-The path to the PDF file
VariableYesText-Variable to store extracted text
CacheNoSelectOnly if validCaching strategy
LanguageNoSelect-Document language

Output variables

VariableTypeDescription
VariableTextThe extracted text from the PDF

Errors

ErrorDescription
File not foundThe file does not exist
Extraction failedCould not extract text from the PDF

Tips

  • Extracts text content from PDF files
  • For scanned PDFs, OCR is used automatically

Related functions

See it working on your own data

Everything documented here ships with the platform – try the document tools free, or go live in 7 days.