Practical guide

How to extract data from PDF files: copy, Excel Power Query, Tabula or a parser

Last materially reviewed 2026-10-03

Quick answerFor a text PDF, try copy and paste, then Excel's Power Query PDF connector. For scans, repeated documents or line items, use an OCR-capable parser. Start with the free options.

First check whether the PDF holds text or a picture of text

To extract data from a PDF, first find out whether it is a text PDF or a scan, because that decides which methods can work at all. Open the file and try to select a single word with your cursor. If the word highlights, the PDF contains real text and the free methods below are worth trying. If nothing selects, the page is an image and you need OCR, which converts pictures of characters into text. Our page on scans and text PDFs goes further.

The first three methods in the table below need no parser subscription. All four are described here from Microsoft's and the tool makers' documentation, not from hands-on use.

Four methods and when each one works

PDF data extraction methods compared
MethodWorks onGood forStops working when
Copy and pasteText PDFsOne value or one small table, onceColumns collapse into one line, or you have more than a few files
Excel Power Query PDF connectorTables in PDFs, per Microsoft's connector pageRepeating the same table import, or a folder of similar PDFsThe data is not laid out as a table, or rows wrap across lines
Tabula (free tool)Text-based PDFs onlyPulling tables to CSV or Excel on your own computerThe PDF is scanned; Tabula's site says it does not handle scans
Parser tool with OCRText PDFs, scans and imagesNamed fields and line items from documents that keep arrivingVolume is too low to justify a subscription

Steps for Excel's Power Query PDF connector

Microsoft lists the PDF connector as generally available in Excel and Power BI, with no prerequisites. The steps in its documentation are:

  1. In Excel's get-data options, choose PDF in the connector list.
  2. Browse to the PDF file and select Open.
  3. In the Navigator window, pick the table or page you want.
  4. Select Load to put it straight into the sheet, or Transform Data to clean it in the Power Query Editor first.

Microsoft documents three practical limits. To import many PDFs at once you use the Folder connector and combine files. Large PDFs can time out, and the suggested fix is to read a few pages at a time with the StartPage and EndPage options. Rows that span several lines may not be recognized, and you may need to repair them with fill-down or grouping steps. Our Power Query first page explains why this free route is worth trying before any paid tool.

Steps for a parser tool, using Parsio as the example

A parser is the method for scans, for documents that arrive every week, and for fields that are not in a table, such as an invoice number in a header. Parsio's help center gives these steps for its AI engine:

  1. Create an inbox, choose the AI-powered PDF parser and select a pre-built model, for example invoices, receipts or bank statements.
  2. Upload files, email them as attachments, or send them through the API.
  3. Let the model identify fields and tables. There is no template to draw.
  4. Export to Google Sheets, a file, a webhook or an automation platform.

Cost per page depends on the engine: 3 credits for the AI parser, 2 for the GPT-powered parser, 1 for the OCR converter, which turns a scan into text and tables without picking out named fields. Parsio's free plan has 30 credits a month. Starter was $29 a month for 100 credits on monthly billing as read on 3 October 2026.

A worked example: a 12-page supplier price list

A retailer receives a 12-page price list as a text PDF every quarter and needs the product code, description and price columns in Excel. Suppose copy and paste merges the columns. The Power Query connector is the right next attempt: the table is regular, the file is text, and the same steps can be repeated each quarter. Cost: nothing beyond Excel.

Now suppose the supplier starts sending a scanned copy. Tabula does not handle scans. In Parsio, the OCR converter would use 12 credits to turn the pages into a table, which fits inside the free plan's 30 credits a month.

The mistake to avoid: paying for a parser to do a one-off job

The mistake to avoid is choosing the method by the tool you have heard of instead of by the document. A single text PDF with one clean table does not need a subscription. A stack of scanned invoices will not yield to copy and paste however long you try. Check text versus scan, count how many documents arrive each month, and decide whether you need a whole table or a few named fields. If the documents are invoices with line items, read the line items page. If you are ready to compare paid tools, see PDF data extraction software.

Sources used for this page

The facts above come from the pages below, read on 3 October 2026. We have not used these products hands-on. Prices and terms change, so confirm them on the vendor's site before you buy.

  1. Power Query PDF connector | Microsoft Learn — Vendor documentation · learn.microsoft.com · checked 2026-10-03
  2. Tabula: Extract Tables from PDFs — Vendor product page · tabula.technology · checked 2026-10-03
  3. OCR Parsing of PDF Files and Images | Parsio Knowledge Base — Vendor documentation · help.parsio.io · checked 2026-10-03
  4. Choosing the Right Parser Type | Parsio Knowledge Base — Vendor documentation · help.parsio.io · checked 2026-10-03
  5. Pricing and Free Trial | Parsio — Vendor pricing page · parsio.io · checked 2026-10-03