From Upload to Source-Linked Evidence

From raw file upload to source-linked findings a reviewer can check: how a document set moves through DataStruct AI.

01

Upload Documents

Drag and drop PDFs (including scans), DOCX, XLSX, CSV, HTML, or plain text files. Upload individually or in bulk (up to 50 at once). DataStruct AI deduplicates automatically by content hash.

PDF (incl. scanned), DOCX, XLSX, CSV, HTML, TXT supported
Bulk upload with batch processing
Automatic deduplication by SHA-256 hash
OCR for scanned documents
02

Automated Processing

A 9-stage pipeline extracts text, detects sections, splits sentences, identifies entities and metrics, creates searchable chunks, and generates vector embeddings.

OCR fallback for scanned documents
Section and heading detection
Sentence-level verbatim storage
Entity extraction (dates, amounts, percentages)
03

Evidence Programmes

Write your requirements in your own words and run them against a named subject — a supplier, a site, a contract. You get one finding per requirement with the candidate evidence attached. In Founding Customer workspaces each finding separates supporting passages from sources cited but not confirmed and from merely related material.

Requirements persist between runs
One finding per requirement, with source links
Supporting / cited-but-unverified / related kept separate
Reruns retain earlier findings
04

Search & Ask

Hybrid search combines keyword matching, semantic vectors, and metadata filters. Ask natural language questions and get evidence-backed answers with citations.

Lexical + semantic + metadata hybrid search
Evidence-grounded answers with page citations
Verification pass flags claims it cannot match to the cited evidence
Insufficient evidence detection
05

Analyse & Extract

Define extraction templates, run compliance rule packs, compare documents side-by-side, and build benchmark groups across your document corpus.

Template-based structured extraction
Declarative rules with deterministic condition types
Side-by-side document comparison
Peer group benchmarking
06

Review & Export

Reviewers confirm, reject, or mark findings for follow-up; each decision is attributed and logged. Extraction datasets export to CSV, XLSX or JSON and document reports to HTML, DOCX or JSON.

Reviewer decisions: confirm / reject / follow-up
Audit log on each decision
Dataset export: CSV, XLSX, JSON
Report export: HTML, DOCX, JSON