Skip to content

Find Text Positions In A PDF

Use PDFReader.extractPageText(pageIndex) to enumerate text-showing operations, then filter their decoded text. Reader pages are zero-based and positions use PDF's bottom-left coordinate system.

import { createMuhammaraWasm } from "@muhammara/wasm";

var muhammara = await createMuhammaraWasm();
var reader = await muhammara.createReaderAsync(pdfFile);
var target = "Text to locate";

try {
  var positions = reader
    .extractPageText(0)
    .filter(function (element) {
      return element.text === target;
    })
    .map(function (element) {
      return {
        x: element.textMatrix[4],
        y: element.textMatrix[5],
        fontSize: element.fontSize,
        fontResource: element.fontResource,
      };
    });
} finally {
  reader.end();
}

Each result is a PDF text-showing operation in direct content-stream order. textMatrix is [a, b, c, d, e, f]; e and f are its origin in page coordinates. The matrix combines explicit text positioning through BT, Tm, Td, TD, TL, T*, ', and " with the active graphics transformation from cm; its first four values retain rotation, scale, or skew. text is the string decoded through the active font: the font's /ToUnicode CMap first, then a simple font's /Encoding and /Differences; codes the font does not map become U+FFFD, as do codes of fonts whose built-in encoding cannot be read, the Symbol and ZapfDingbats standard fonts, embedded Type 1 fonts without /Encoding, and Type 3 fonts, unless /ToUnicode or /Differences maps them, and every byte of text shown without a font or with a font that cannot be resolved or read. content keeps the raw character codes, for example two-byte glyph IDs for text written with a composite font. A phrase split across operations does not match as a whole. The extractor does not calculate glyph bounds or glyph-driven matrix advances, or descend into Form XObjects such as appended content created by Recipe.editPage(). Adjacent text-showing operations without an explicit positioning operator retain the same matrix. This is not a general visual full-text search or glyph-bounds API.

The browser example's Find text tab runs this search on an uploaded PDF, or on a built-in sample when none is chosen, and highlights each match. It estimates match widths with Recipe's bundled font, so highlights are exact for the sample and approximate for PDFs that use other fonts.

Bound the work on untrusted input

Both extractors accept an optional limits object. Fields you omit keep the built-in default, and values above it are clamped down, so a caller can tighten the budget but never raise it past the ceiling:

var reader = await muhammara.createReaderAsync(pdfFile);
var elements = reader.extractPageText(0, {
  maxElements: 500,
  maxTextBytes: 64 * 1024,
});
reader.end();
Field Default and ceiling
maxElements 100000
maxOperands 1024
maxTextBytes 16777216 (16 MiB)
maxParsedObjects 1000000

A page that exceeds the budget throws rather than returning partial results.

Independent of limits, an extraction reads at most 64 MiB of a page's decoded content, inline image data included, and throws the same error past it. A compressed content stream can expand far beyond its file size, so this bounds the work one small PDF can cause.

Decoding text also reads the page's fonts. The PDF objects it reads count against maxParsedObjects on their own, and exceeding it throws the same error. Fonts are cached per reader, but each call still counts a cached font's objects and CMap bytes, so results never depend on earlier calls. A /ToUnicode CMap larger than 4 MiB, or beyond 32 MiB of CMaps in one call, is ignored and that font decodes through its /Encoding instead. A font that cannot be read decodes to U+FFFD rather than failing the call. When you only need positions or raw content, pass { decodeText: false } as the third argument to skip reading fonts; the elements then have no text.

Detect page marks without reading text

PDFReader.extractPageContentItems(pageIndex, limits?) reports every direct content-stream operation that puts a mark on the page, which is a cheaper way to answer "is this page blank?" than extracting text:

var reader = await muhammara.createReaderAsync(pdfFile);
var isBlank = reader.extractPageContentItems(0).length === 0;
reader.end();

Each item is { type, operation }, where type is one of ePDFPageContentItemText, ePDFPageContentItemPath, ePDFPageContentItemXObject, or ePDFPageContentItemShading — read from the loaded runtime, as in muhammara.ePDFPageContentItemPath — and operation is the PDF operator that produced it. An inline image reports as ePDFPageContentItemXObject with operation BI; its binary payload is skipped rather than parsed, so the bytes cannot invent extra items. Text drawn in an invisible rendering mode (Tr 3 or Tr 7) is excluded; white-on-white text is included, because it is still a page mark. limits.maxTextBytes is accepted for signature parity with extractPageText but has no effect here.