HomeVideos

KDD 2026 - Benchmarking Table Extraction from Heterogeneous Scientific PDF Documents

Now Playing

KDD 2026 - Benchmarking Table Extraction from Heterogeneous Scientific PDF Documents

Transcript

50 segments

0:00

Hi everyone, I'm Marin Šoštarić, and

0:02

today I will present our paper on

0:04

benchmarking table extraction. Tables in

0:06

PDF contain valuable data that we want

0:09

to extract and exploit later. These

0:11

tables are visually structured for

0:13

humans, but PDF has no machine-readable

0:15

structure, which makes table extraction

0:17

harder.

0:18

We define two subtasks. In TD, we locate

0:21

the table bounding box, then in TSR plus

0:24

TCR, we extract its structure and

0:26

content. Table extraction is the overall

0:29

end-to-end task. We ask ourselves, how

0:31

do we best evaluate a table extraction

0:33

model end-to-end? Prior work tackles the

0:36

two subtasks independently. In TD, the

0:38

input is a page image, and we use

0:41

classical retrieval metrics. In TSR, the

0:44

input is a perfectly cropped table.

0:46

The two evaluations are done in

0:48

isolation with no relationship between

0:50

them.

0:52

However, table detection feeds the table

0:55

structure recognition model, so errors

0:57

in the first stage propagate to the

1:00

second.

1:03

This is why we propose a new end-to-end

1:06

metric that accounts for the dependency

1:08

between the two tasks.

1:12

We benchmark 15 models on four datasets,

1:15

introduce end-to-end metrics.

1:17

We also study inference cost and model

1:19

calibration for subset of models. Two of

1:22

the datasets are new, one automatically

1:24

generated from our craft sources, one

1:26

manually annotated. They also contain

1:28

negative samples, which balances

1:30

precision to better reflect real-world

1:32

performance.

1:33

We benchmark 15 models ranging from

1:35

heuristic, computer vision, general, and

1:38

visual language models to commercial

1:39

tools.

1:40

Model performance depends strongly on

1:43

document layout.

1:45

The most expensive models are not

1:47

necessarily the best.

1:49

Thank you for your attention. Find more

1:51

details in the paper.

Interactive Summary

Marin Šoštarić presents a research paper on benchmarking table extraction from PDF documents. The study addresses the limitations of evaluating table detection and structure recognition in isolation, proposing a new end-to-end metric that accounts for error propagation between these stages. The evaluation encompasses 15 diverse models across four datasets, including two newly introduced ones, and provides insights into inference costs and the impact of document layout on performance.

Suggested questions

3 ready-made prompts