Skip to main content
Skip to article
TablefoldExcel you can still check

checklist

How do you turn PDF tables that span pages into one continuous sheet?

Append each page's rows beneath the last, strip the header block that repeats on every page after the first, and repair any row the page break cut in half — then prove the merge with a row count and a total before anyone uses the sheet.

Why multi-page tables arrive shattered

A paginated report restarts its table on every page: header row again, sometimes carried-forward subtotals, occasionally a footer with page numbers sitting inside the table region. Converted naively, ten pages become ten fragments stitched with repeated headers that masquerade as data rows — and a subtotal row from page six lands in the middle of July.

Before working the checklist, one status note: you can open these pages and sign in to the workspace today, but that does not mean external platform accounts are connected, and it does not mean the product has formally launched. Everything below runs on the PDFs you supply.

What you need before you merge

The source PDF with its page boundaries known, and the grand total printed somewhere in the report. The total is the proof that the merged sheet kept everything and added nothing.

A decision on the dedupe key before merging starts: which combination of columns makes a row genuinely identical to another. Choosing the key after seeing duplicates is how legitimate repeated values get deleted along with the noise.

Step 1: Inventory what repeats on every page

Open the PDF and note what recurs per page: the header row, running subtotals, footers, page numbers. Each recurring element needs a removal rule of its own, and knowing them up front beats discovering them as anomalies mid-merge.

Step 2: Convert page by page and append in order

Convert each page's table and append beneath the previous one in reading order. Page-at-a-time keeps every fragment attributable to its source page, which matters the moment something needs tracing back.

Step 3: Remove repeated headers by rule, not by eye

Delete any row whose key fields match the header text exactly. Doing it by eye across forty pages invites one missed header, and a missed header becomes a text value inside a numeric column, breaking sums quietly.

Step 4: Repair rows split at the page boundary

A row cut by a page break arrives as two fragments: the top half ends one page, the bottom half opens the next. Rejoin fragments whose fields are incomplete on their own — a date with no amount above, an amount with no date below — and verify each join against the source page pair.

Step 5: De-duplicate with the key you chose

Remove rows identical across the full dedupe key. Identical transactions on the same day are real data, so resist broadening the key to make deduplication easier; a narrower key leaves duplicates in place, which the totals check will surface anyway.

Step 6: Prove the merge with arithmetic

Count rows in the sheet and compare against the sum of rows on the source pages, accounting for removed headers and rejoined splits. Then total the sheet against the report's printed grand total. Both checks passing is what turns ten fragments into one sheet you can hand over.

Verification

The merge holds when no header row survives past the first occurrence, the row count matches the source arithmetic, and the grand total agrees with the printed figure. A colleague should be unable to tell from the sheet alone where one page ended and the next began.

Limits worth stating plainly

Tables are not cell-perfect. Check every number against the source page before you rely on it. A whole multi-page report converted in one pass is not reliable today. Rows from different periods can land in one cell and a column can be dropped. Work one table at a time. Audit-style statements with dollar signs, em-dash blanks, and parenthesised negatives are not handled. Data rows can be lost. Scanned PDFs with no text layer are refused rather than guessed at. There is no OCR path you can rely on here today. A document with no printed table is not something to convert. If you are handed a sheet built out of running prose, discard it and tell us.

Tablefold detects the grid on each page and flags spans it was unsure about — including the boundary regions where splits hide — but deciding which repeated block is a header and which fragment pairs with which stays your call, confirmed against the source pages.

What Tablefold does in this workflow

Tablefold converts the paginated report and marks low-confidence cells, so the checklist runs against flagged material rather than raw guesses. Boundary regions between pages get the same visibility as anywhere else in the table.

The merge rules remain yours to set and apply: what repeats, what the dedupe key is, and when the arithmetic proves the job done. The product's contribution is a consistent grid per page and honest flags where detection strained.

FAQ

Questions this guide is for

Should I remove the repeated header rows or keep one?

Keep exactly one, at the top. Every extra header row puts text into a numeric column, which breaks sums and sorts downstream in ways that surface far from the cause.

Two rows look identical. Are they duplicates?

Only if they match across the full dedupe key you defined before merging. Same-day identical transactions are common and real, so delete on the agreed key alone and let the totals check catch anything left behind.

What if the table continues without repeating the header?

Then there are fewer rows to remove, but the page-boundary split risk stays the same. Check the first data row of each page for completeness against the last row of the previous page.

Start in the workspace

Merge a paginated report into one sheet

Sign in or create an account and you return to the Tablefold conversation. Upload the multi-page report, convert the tables, and use the flagged grid to run the merge checklist against the source.

Tablefold

Signing in and billing happen in the conversation. This page uses PostHog for product analytics (anonymous, optional). See Privacy.