Thin wrapper around a PDFBox PDDocument for a single loaded PDF, exposing page access, AcroForm field reading and writing, and document metadata. Created by FileManager.parsePDF(...) or directly by PdfManager when generating thumbnails and other derived output. The wrapped PDDocument must be closed via close() once the caller is finished with it.


Properties

PropertyReturnsDescription
documentCatalogPDDocumentCatalogThe PDF document's root catalog dictionary (the /Root entry), giving access to top-level structures such as the AcroForm and the page tree. Never null.
documentInformationPDDocumentInformationThe document information dictionary (the /Info entry), holding metadata such as title, author and creation date. Never null.
encryptedbooleanWhether this PDF was loaded from an encrypted source document.
encryptionPDEncryptionThe encryption parameters recorded on this document, if it is encrypted. PDF supports pluggable encryption handlers, but PDFBox currently only implements the standard security handler.
fieldsList<PDField>The document's root AcroForm fields. A field may itself be a container of further fields (a non-terminal field) or a leaf value field (a terminal field), and the roots may include both kinds mixed together; call getChildren on a non-terminal field to walk further down the tree.
numberOfPagesintThe number of pages in this PDF document.
pagesPDPageTreeThe pages of this document, in document order.
pDFRendererPDFRendererThe PDFBox renderer for this document, used to rasterise pages to images. Lazily created and cached on the first call.

Methods

getDocumentCatalog() · isEncrypted() · getEncryption() · setEncryptionDictionary(PDEncryption encryption) · getPages() · getPage(int pageIndex) · getNumberOfPages() · setField(String fieldName, String value) · setFields(Map<String,String> params) · fieldExists(String fieldName) · getFields() · getDocumentInformation() · save() · close() · getPDFRenderer()

getDocumentCatalog()

Returns: PDDocumentCatalog

The PDF document's root catalog dictionary (the /Root entry), giving access to top-level structures such as the AcroForm and the page tree. Never null.

isEncrypted()

Returns: boolean

Whether this PDF was loaded from an encrypted source document.

getEncryption()

Returns: PDEncryption

The encryption parameters recorded on this document, if it is encrypted. PDF supports pluggable encryption handlers, but PDFBox currently only implements the standard security handler.

setEncryptionDictionary(PDEncryption encryption)

Returns: void

Replaces the encryption parameters recorded on this document.

ParameterDescription
encryptionthe encryption dictionary to apply, typically a standard security handler instance

getPages()

Returns: PDPageTree

The pages of this document, in document order.

getPage(int pageIndex)

Returns: PDPage

Returns the page at the given index.

ParameterDescription
pageIndexthe page index

getNumberOfPages()

Returns: int

The number of pages in this PDF document.

setField(String fieldName, String value)

Returns: void

Attempts to set a value of a field

ParameterDescription
fieldNamethe field name
valuethe new field value.

setFields(Map<String,String> params)

Returns: void

Sets multiple AcroForm field values in one call, keyed by field name.

ParameterDescription
paramsa map of field values, keyed on the field name; ignored if null or empty

fieldExists(String fieldName)

Returns: boolean

Checks if field exists

ParameterDescription
fieldNamethe field name

getFields()

Returns: List<PDField>

The document's root AcroForm fields. A field may itself be a container of further fields (a non-terminal field) or a leaf value field (a terminal field), and the roots may include both kinds mixed together; call getChildren on a non-terminal field to walk further down the tree.

getDocumentInformation()

Returns: PDDocumentInformation

The document information dictionary (the /Info entry), holding metadata such as title, author and creation date. Never null.

save()

Returns: String

Saves the current in-memory state of the wrapped PDDocument and stores the resulting bytes in the blob store, splitting and hashing the output the same way as other stored content.

close()

Returns: void

This will close the underlying PDF Document.

getPDFRenderer()

Returns: PDFRenderer

The PDFBox renderer for this document, used to rasterise pages to images. Lazily created and cached on the first call.

To get full access to the Kademi Hub existing customers can login here, or new customers can register here.