Add documents
Knovas searches the text you send it, so the way you prepare that text affects search quality.
How uploading works
- Prepare the document as plain text.
- Start an upload and say how many parts you will send. You get back an upload key.
- Send each part with that key. When the last part arrives, Knovas processes the document in the background.
Step 1: Prepare the text
Knovas does not open PDF or Word files for you. Your software extracts the text first. For the best results:
- Use simple Markdown. Mark headings with
#,##and###. Knovas uses the headings to understand the structure. There is no separate “chapter” field. - Remove repeated clutter such as page headers, footers and page numbers.
- Split at natural breaks, such as paragraphs or sections. About 500 characters per part works well.
Our free Python library knovas-extract (pip install knovas-extract) turns PDF, Word, HTML, RTF, email and text files into clean text and tables in exactly the format Knovas expects.
Step 2: Start the upload
curl -X POST https://api.knovas.ch:8443/secured/init_document_transmission \
--cert client_cert.pem --key client_key.pem --cacert ca_root_cert.pem \
-H "Content-Type: application/json" \
-d '{
"identifier": "reports/2025/q3.pdf",
"part_count": 3,
"title": "Q3 Financial Report 2025",
"description": "Quarterly results, revenue by region and outlook.",
"path": "/Reports/2025/q3.pdf"
}'import requests
BASE = "https://api.knovas.ch:8443"
AUTH = dict(
cert=("client_cert.pem", "client_key.pem"), # your certificate and private key
verify="ca_root_cert.pem", # Knovas' certificate
timeout=60,
)
r = requests.post(f"{BASE}/secured/init_document_transmission", **AUTH, json={
"identifier": "reports/2025/q3.pdf",
"part_count": 3,
"title": "Q3 Financial Report 2025",
"description": "Quarterly results, revenue by region and outlook.",
"path": "/Reports/2025/q3.pdf",
})
r.raise_for_status()
key = r.json()["transmission_key_id"]| Field | Required? | What it does |
|---|---|---|
identifier | Yes | Your own name for the document (up to 1000 characters). Shown as pointer in results. |
part_count | Yes | How many parts you will send (1 to 10,000). |
title | No, but strongly recommended | The strongest signal for search. Up to 500 characters. |
description | No | A short description in your own words. Up to 2000 characters. |
path | No | Folder and file name. Words in it help search (e.g. a folder called Contracts). |
access_groups | No | Who may see the document, e.g. ["HR"]. See Access control. |
The answer (status 201) contains transmission_key_id. You need it in the next step.
Step 3: Send the parts
Send parts numbered 0, 1, 2 … up to part_count - 1.
curl -X POST https://api.knovas.ch:8443/secured/transmit_document_part \
--cert client_cert.pem --key client_key.pem --cacert ca_root_cert.pem \
-H "Content-Type: application/json" \
-d '{
"key": "<transmission_key_id>",
"part_number": 0,
"snippet": "# Q3 Financial Report\n\n## Highlights\n\nRevenue grew by 8% ...",
"page_number": 1
}'import requests
BASE = "https://api.knovas.ch:8443"
AUTH = dict(
cert=("client_cert.pem", "client_key.pem"), # your certificate and private key
verify="ca_root_cert.pem", # Knovas' certificate
timeout=60,
)
parts = [
"# Q3 Financial Report\n\n## Highlights\n\nRevenue grew by 8% ...",
"## Revenue by region\n\nEMEA grew fastest ...",
"## Outlook\n\nWe expect ...",
]
for number, text in enumerate(parts):
r = requests.post(f"{BASE}/secured/transmit_document_part", **AUTH, json={
"key": key, # from step 2
"part_number": number,
"snippet": text,
})
r.raise_for_status()
print(r.json()["transmission_complete"]) # True after the last part| Field | Required? | What it does |
|---|---|---|
key | Yes | The transmission_key_id from step 2. |
part_number | Yes | The number of this part, starting at 0. |
snippet | Yes | The text of this part. |
page_number | No | The page this part starts on. Returned in search results. |
sentence_number | No | The number of the first sentence in this part. Returned in search results. |
tables | No | Tables found in this part. See Tables. |
The answer for the last part says "transmission_complete": true. Knovas then processes the document in the background, which usually takes seconds to a minute. After that it shows up in searches.
Updating a document
Upload it again with the same identifier. The new version replaces the old one. If the same content is uploaded under two different identifiers, Knovas stores it only once, and both identifiers keep working.
Page and sentence numbers
If you send page and sentence numbers, search results can point to the exact place in the document, so your users can jump straight to it.
- Page:
page_numberis the page a part starts on. If a part crosses into the next page, put a page-break character (\f) at the break. Most PDF tools add it automatically. Or simply send one page per part. - Sentence:
sentence_numberis the number of the first sentence in the part; Knovas counts on from there. If you send no sentence numbers at all, Knovas numbers the sentences for you. Either number all parts or none.
Tables
Send real tables as structured data next to the text. Knovas then keeps rows and columns together, so a question like “APAC revenue” finds the right cell.
curl -X POST https://api.knovas.ch:8443/secured/transmit_document_part \
--cert client_cert.pem --key client_key.pem --cacert ca_root_cert.pem \
-H "Content-Type: application/json" \
-d '{
"key": "<transmission_key_id>",
"part_number": 1,
"snippet": "## Revenue by region\n\nSee the table below.",
"tables": [
{
"client_table_hint": "revenue-by-region",
"title": "Revenue by region (Q3 2025)",
"headers": ["Region", "Revenue", "Change"],
"rows": [["EMEA", "12.4M", "+8%"], ["APAC", "9.1M", "+14%"]]
}
]
}'import requests
BASE = "https://api.knovas.ch:8443"
AUTH = dict(
cert=("client_cert.pem", "client_key.pem"), # your certificate and private key
verify="ca_root_cert.pem", # Knovas' certificate
timeout=60,
)
r = requests.post(f"{BASE}/secured/transmit_document_part", **AUTH, json={
"key": key,
"part_number": 1,
"snippet": "## Revenue by region\n\nSee the table below.",
"tables": [{
"client_table_hint": "revenue-by-region",
"title": "Revenue by region (Q3 2025)",
"headers": ["Region", "Revenue", "Change"],
"rows": [["EMEA", "12.4M", "+8%"], ["APAC", "9.1M", "+14%"]],
}],
})
r.raise_for_status()Each row must have exactly as many cells as there are headers. Limits: 50 tables per part, 64 columns, 5000 rows and 1024 characters per cell.
Upload checklist
- Always send a meaningful
title. Documents without titles are the most common cause of weak results. - Add a
pathif your documents live in folders. - Send clean text with headings, not a raw copy of the page layout.
- Use one identifier per document, and keep it stable over time.
- If you use access control, send
access_groupswith every upload.
Questions? Write to contact@knovas.ch.
This page describes Knovas 1.3.0. Last updated .