Document search
Ask questions across millions of pages — PDF, Word, spreadsheets and scanned images via OCR. Every answer cites file and page.
ON-PREMISE AI PLATFORM
Peki deploys local language models with a hybrid RAG engine — vector tables, SQL and a knowledge graph — over your PDFs, Word files and scans. Answers with citations, on your hardware, with zero internet dependency.
TODAY · 09:41
List the termination notice periods in our active vendor contracts signed after 2022.
PEKI-33B · LOCAL
Found 7 active contracts signed after Jan 2022. Notice periods:
+ 4 more · 47 contracts scanned · every claim cited
CAPABILITIES
Ask questions across millions of pages — PDF, Word, spreadsheets and scanned images via OCR. Every answer cites file and page.
Vector similarity for meaning, SQL for exact fields and dates, a knowledge graph for relationships — merged into one grounded answer.
Local models tuned on your terminology, formats and policies. From 7B on a single workstation to 70B on a GPU node.
Multi-step agents for complex tasks: contract review, compliance checks, report assembly — each step logged and auditable.
ARCHITECTURE
Documents are parsed, OCR’d and chunked once — then indexed three ways so every question finds the right kind of evidence.
RUNS ENTIRELY INSIDE YOUR NETWORK
PDF · DOCX · XLSX
TIFF · PNG scans
parse · OCR
chunk · embed
Vector tablesVectorANN
SQLSQLEXACT
Knowledge graphGraphLINKS
rank · merge
deduplicate
answers + citations
agents · tools
Vector — finds passages by meaning, even when wording differs across documents.
SQL — exact filters over extracted fields: dates, amounts, parties, statuses.
Graph — entities and relations, so amendments, parties and clauses stay connected.
SECURITY
There is no cloud fallback and no telemetry endpoint to disable — the software has no route out of your network to begin with.
YOUR NETWORK
file server · DMS
vector · SQL · graph
GPU node / appliance
browser · LAN only
AIR GAP
no route out
DEPLOYMENT
01 · DAYS 0–3
We audit your document estate, tasks and hardware. You get a sizing plan — GPU node or shipped appliance.
02 · WEEK 1
Install on your network — no inbound or outbound internet required. Integration with AD/LDAP and file shares.
03 · WEEKS 2–3
Ingestion pipeline parses, OCRs and indexes your documents into vector tables, SQL and the knowledge graph.
04 · WEEK 4
We tune chat models on your terminology and configure agents for your workflows. Your team goes live.
KNOWLEDGE BASES
Curated, versioned and citable corpora that ship with your deployment — updated through the same signed offline bundles.
Statutes, case law, contract doctrine and regulatory texts.
2.1M SOURCES
Clinical guidelines, drug interactions, coding and protocols.
870K SOURCES
Accounting standards, tax codes and reporting frameworks.
640K SOURCES
Industry norms, standards and technical documentation.
1.3M SOURCES
We build and maintain a corpus for your domain, to order.
BUILT TO ORDER →
CONTACT
No sales script. Describe your documents and tasks — we’ll tell you what hardware it needs and what a pilot looks like.