Lexyra AI analyzes PDFs at the rendering layer to detect hidden text and images before they reach AI agents. Its deep parsing engine reconstructs their graphical context, hierarchy, and deep document structure for use in agentic systems, RAG pipelines, and enterprise workflows.
Every agentic AI system that ingests PDFs is exposed to indirect prompt injection risk. Hidden instructions can pass through conventional parsers and downstream guardrails before the agent ever reads them. The attack surface begins inside the document itself.
In agentic AI, a PDF is no longer a passive content. Extracted text and images can become instructions executed by an agent. Hidden layers, white-on-white text, microscopic fonts, and off-page elements can silently trigger data leakage, corrupt reasoning, or hijack an autonomous workflow.
In 2025, a Nikkei Asia investigation found hidden prompts in at least 17 arXiv preprints whose authors were affiliated with 14 academic institutions across eight countries. The prompts were intended to influence AI-assisted peer review. These were real research manuscripts, not lab-generated proofs of concept. In 2026, the Center for Internet Security warned that malicious instructions hidden in documents can lead to data theft, unauthorized access, and disrupted operations.
Most AI guardrails analyze prompts, extracted text, or model outputs. Existing PDF-level approaches remain limited and unreliable, even on basic hidden-content cases. Lexyra AI inspects documents at the rendering layer before ingestion, where concealed content can be detected at its source.
PDFs must also be inspected on the way out. A compromised agentic workflow can generate an apparently legitimate document containing hidden instructions for downstream AI systems or concealed sensitive data. Lexyra AI scans generated PDFs before they leave the pipeline, preventing propagation and covert exfiltration.
Lexyra AI analyzes PDFs at the rendering layer, where concealed content becomes structurally and visually inspectable. Its engine detects suspicious hidden text and images before they reach AI agents, and validates generated PDFs before they leave the workflow.
Majority of concealed content techniques are invisible or easily overlooked during normal document review. All native PDF text parsers will extract the hidden content without realizing that it was invisible to the user, allowing it to enter an AI pipeline as trusted input. Lexyra AI detects these concealment patterns through rendering-aware analysis. Below are eight representative examples from the 20+ techniques mapped by our engine.
Text rendered in the same color as the page background. Invisible during normal viewing, yet extractable by many PDF parsers and potentially ingested as trusted input by downstream AI systems.
Microscopic text that is unreadable during normal viewing but will generally be extracted and passed to downstream AI systems.
Text positioned outside the visible page area, such as beyond the page boundaries or CropBox, while remaining present in the PDF content stream.
Text visually hidden beneath opaque shapes or drawings while remaining present in the PDF content stream. Lexyra AI identifies text that remains extractable despite being occluded during normal viewing.
Text encoded in the PDF with an invisible rendering mode. It is not displayed during normal viewing but is extracted by PDF parsers and enter downstream AI pipelines as trusted input.
Text assigned zero or near-zero opacity values. Invisible or easily overlooked during normal viewing, it will be extracted by PDF parsers and enter downstream AI pipelines as trusted input.
Text clipped outside the visible rendering region while remaining encoded in the PDF content stream. It will be extracted by PDF parsers and enter downstream AI pipelines without appearing during normal viewing.
Text rendered at an unreadable size through graphical scaling, even when its declared font size appears normal. It remains extractable by PDF parsers and enter downstream AI pipelines without attracting attention during normal viewing.
Rule-based filters can detect obvious signatures, but indirect prompt injections do not need to look malicious. Attackers can use ordinary language, contextual instructions, and social engineering techniques that evade simple pattern matching. A syntactically benign sentence may still be designed to manipulate an agentβs behavior.
LLM guards remain valuable, but they can only assess the information they receive. When a PDF is reduced
to extracted text, critical evidence may already be missing: was the sentence visible to the user or
not?
Lexyra AI adds a source-layer control before extracted PDF content is trusted or exposed to
agents. Rendering-aware analysis identifies concealed elements early, while downstream guardrails continue
to protect the rest of the AI pipeline.
PDFs encode visual appearance, not a reliable machine-readable representation of document meaning. Even AI-powered
parsers still lose critical relationships: hierarchy, reading order, nested lists, complex tables, captions,
notes, links, and graphical context.
When incomplete or incorrectly rebuilt structure enters an AI pipeline,
every downstream task becomes less reliable, including retrieval, reasoning, generation, and automation.
Garbage in, garbage out, silently and at scale.
When invisible text is embedded alongside genuine document content, it can alter vector embeddings and manipulate semantic-search rankings before any model is queried. Without rendering-aware inspection, a text-only RAG pipeline treats concealed content as trusted data because the rendering context has already been lost.
Complex tables become disconnected rows of text. Headings lose their hierarchy. Nested lists, captions,
notes, and cross-references detach from the elements they describe. Multi-column layouts collapse into
unreliable reading sequences.
By the time extracted content reaches an AI system, critical relationships
may already be lost. Retrieval becomes less accurate, reasoning loses context, and every downstream task
becomes less reliable.
Even AI-powered parsers still lose critical document relationships: nested hierarchy, reading order, complex
table semantics, captions, notes, links, and graphical context. Once this structure is missing from the
extracted data, downstream AI systems cannot reliably recover it.
The failure starts at the source,
long before the model generates an answer.
Lexyra AI rebuilds the semantic and visual architecture of complex PDFs, not just their text. Its deep extraction engine preserves the relationships that AI pipelines need, and that extraction tools still frequently lose.
Left: bounding box overlay on the original PDF page, showing detected elements with their spatial coordinates. Right: structured JSON output with full hierarchy, table reconstruction, and hidden content detection results including risk scores.
PDF parsers extract content. AI-powered extraction tools recover part of the document structure. LLM guardrails inspect text, prompts, and model behavior. One engine, two controls: rendering-aware threat detection and deep document extraction before PDF content enters your AI pipeline.
Measured on native PDFs: hidden-threat detection in under 100 ms per page and deep extraction in 600β900 ms per page.
Current scope: Lexyra AI is optimized for native PDFs. Scanned-document processing with vision and OCR is part of the product roadmap.
Get notified when beta testing opens, request access to our private interactive demos, book an online demo, apply to join our design partner program, or contact us to discuss our pre-seed round.
Thank you. We will reply to as soon as possible.