PDF is a container, not a single kind of research material. One file may contain selectable text, a scanned page image, tables, photographs, marginal annotations, or a mixture of all five. A defensible PDF workflow begins by identifying what is actually present on the page and choosing the coding method that preserves its meaning.

First, test the text layer
Try selecting and copying a short passage into a plain text editor. If the copied wording is accurate and in the correct order, text coding and search are likely appropriate. If characters are missing, words are rearranged, columns merge, or the copied passage is nonsense, the apparent text layer may not be reliable enough for text-based analysis.
OCR can make a scanned document searchable, but it does not guarantee accuracy. Names, non-Latin scripts, hyphenated lines, tables, and older print often produce errors. Treat OCR as a convenience that must be checked against the visible document, not as the source of truth.
Choose the right evidence anchor
Use text selection when the wording itself is the analytical object: a policy claim, interview quotation, or document statement. Use a visual region when the evidence includes layout, image, diagram, handwriting, table position, or a scanned passage without a reliable text layer. A PDF can contain both forms of evidence, and a single project can use both without creating a separate workflow.
Do not lose the page context
A sentence in a report may have a different status when read with its heading, footnote, table, or surrounding argument. Keep enough context in the selected evidence or memo to show what the segment is doing. For long documents, record page numbers, section titles, and whether an excerpt is a quotation, summary, recommendation, or metadata label.
Handling tables and figures
Tables and figures deserve their own decision. If you are studying the values, labels, or visual arrangement, code the relevant region and describe what is visible. If you are treating a table as background information, a memo or document-level note may be more appropriate than coding every cell.
Questions researchers often ask
Can I search a scanned PDF?
Only after OCR has created a sufficiently accurate text layer. Check a sample of meaningful passages before relying on search results.
Can I code a figure and the paragraph that explains it?
Yes. Code the visual region and the relevant explanatory text separately, then use a memo or shared code to preserve their relationship.