Skip to Content
GuidesDocument Processing

Document Processing

Eneo’s document processing capabilities allow you to create AI-powered knowledge bases from your documents and websites. This guide covers uploading documents, web crawling, and how retrieval is configured.

Overview

Eneo can process and extract knowledge from:

  • Documents: PDF, Word (.docx), PowerPoint (.pptx), Excel (.xlsx, .xls), CSV, JSON, XML
  • Websites: Automated crawling and content extraction, optionally including linked documents
  • Text files: Plain text, Markdown

All content is:

  1. Extracted from source format
  2. Chunked into optimal segments
  3. Embedded as vectors for semantic search
  4. Published as one searchable version in PostgreSQL with pgvector
  5. Retrieved when relevant to queries

Updating a document does not remove the working knowledge first. Eneo publishes the replacement text and all vectors together. If extraction or embedding fails, search continues to use the previous complete version. Uploading the same content again with the same embedding model reuses the current version without another embedding call. Changing the model builds compatible vectors before the replacement becomes searchable.

This is separate from file-byte placement. PostgreSQL remains the complete default; compatible object storage is optional and does not replace pgvector.


Uploading Documents

Documents live in collections, and collections belong to a space. Every collection is bound to one embedding model, chosen from the models enabled in the space when the collection is created.

Create a Space

  1. Log in to your Eneo instance
  2. Click Spaces in the sidebar and create a space, or open your personal space
  3. Make sure at least one embedding model is enabled for the space; the collection dialog refuses to continue otherwise

Create a Collection

  1. Open the space and click the Knowledge tab
  2. Under Collections, click Create collection
  3. Give it a name and pick the embedding model

Upload Files

  1. Open the collection
  2. Drag files onto the page or click Upload files

Each file becomes one background job. The job list shows the job as queued, in progress, and finally complete or failed; the frontend polls job status automatically, so there is no need to refresh.

Supported Formats

FormatExtensionNotes
PDF.pdfText-based PDFs; tables are converted to Markdown. OCR is not supported
Word.docxModern Word documents. Legacy .doc is rejected with a “save as .docx” error
PowerPoint.pptxExtracts text from slides. Legacy .ppt is rejected
Excel.xlsx, .xlsEvery sheet is serialised row by row; merged cells are forward-filled
CSV.csvRead as text
JSON, XML.json, .xmlRead as text
Text.txt, .mdUTF-8 first, then Windows-1252 as a fallback

The file type is detected from the content, not the extension. Encrypted PDFs, corrupt archives, and files without any extractable text fail with a specific failure code on the job (encrypted, corrupt, no_extractable_text, unsupported_format).

File Size Limits

Administrators with the Storage permission set the knowledge-file limit in Admin > File storage. Changes apply without restarting backend or worker. Operators keep a deployment safety ceiling for PostgreSQL rows or the optional object-store transport; Eneo uses the lower effective limit and shows which setting constrains it.

See Choose Content Storage for the exact responsibility split.


Web Crawling

A website source crawls a site and stores each page as a document in the space, bound to the embedding model chosen when the website is created (the space’s default embedding model when none is selected).

Add a Website

  1. In your space, open the Knowledge tab and go to Websites
  2. Click Connect website
  3. Fill in:
    • URL (required): where indexing starts, including https://, or the full URL of a sitemap.xml for sitemap crawls
    • Display name: optional, defaults to the URL
    • Crawl type: Basic crawl or Sitemap based crawl
    • Download and analyse compatible files: also fetch linked PDFs, Office documents, CSV, JSON, XML, and text files (basic crawls only)
    • Automatic updates: Never, Every day, Every other day, or Every week
    • HTTP Basic Authentication: optional username and password for protected sites; credentials are encrypted and used only for this website
    • Embedding model: shown only when creating; defaults to the space’s default embedding model
  4. Click Create website

The scheduler checks every hour on the hour (UTC): a daily website is queued once at least 24 hours have passed since its last crawl, an every-other-day website after 48 hours, and a weekly website only on Fridays (UTC) once seven days have passed. Use Start crawl on the website to run a crawl manually at any time.

Starts at the URL and follows links within the site until the tenant’s page limit or time limit is reached. robots.txt is respected by default and request rate is auto-throttled. This is the only mode that can download linked files.

Monitor Progress

Each crawl is recorded as a crawl run with its page count and status, visible on the website page and over the WebSocket. If a website fails ten crawls in a row it is automatically disabled; change its update interval to re-enable it.

Crawler Limits

Crawl limits are not set per website. They are deployment defaults from env_backend.env that a system administrator can override per tenant.

SettingDefaultPurpose
CRAWL_MAX_LENGTH36000 s (10 hours)Maximum duration of one crawl
CLOSESPIDER_ITEMCOUNT20000Maximum pages per crawl
DOWNLOAD_MAX_SIZE10485760 (10 MB)Maximum size of a downloaded file
OBEY_ROBOTStrueRespect robots.txt
AUTOTHROTTLE_ENABLEDtrueSlow down based on the target server’s response time
CRAWL_PAGE_BATCH_SIZE100Pages committed per database write
CRAWL_EMBEDDING_CONCURRENCY3Parallel embedding calls per crawl
TENANT_WORKER_CONCURRENCY_LIMIT4Concurrent crawl jobs per tenant
CRAWL_FEEDER_ENABLEDtrueRelease queued crawls in batches
USING_CRAWLtrueEnable crawling for the deployment

Per-tenant overrides are stored on the tenant and read by the worker at the start of every crawl:

# Requires the super API key curl -X PUT https://your-eneo-instance.com/api/v1/sysadmin/tenants/<tenant-id>/crawler-settings \ -H "X-API-Key: $ENEO_SUPER_API_KEY" \ -H "Content-Type: application/json" \ -d '{"closespider_itemcount": 5000, "crawl_max_length": 7200}'

GET on the same path returns the effective settings, and DELETE reverts the tenant to the deployment defaults. Every setting is range-checked; the response lists the valid range when a value is rejected.

Web Crawling Best Practices

  1. Start small: test with a section of the site before crawling all of it
  2. Prefer sitemaps for large sites where the sitemap is maintained
  3. Check robots.txt: the crawler honours it unless the tenant disables obey_robots
  4. Use specific URLs: point a basic crawl at a documentation section, not the root domain

How Document Processing Works

1. Extraction

Content is extracted from source formats:

  • PDF: pdfplumber, with tables rendered as Markdown
  • Word: docx2python
  • PowerPoint: python-pptx
  • Excel: pandas with the calamine engine
  • Web: Scrapy spiders, HTML parsed with BeautifulSoup and converted with html2text

2. Chunking

Text is split with LangChain’s RecursiveCharacterTextSplitter. Chunk size is measured in tokens, not characters:

  • Chunk size: 200 tokens
  • Overlap: 40 tokens

Tokens are counted with LiteLLM’s model-aware tokenizer, falling back to a four-characters-per-token estimate if counting fails.

The splitter reads CHUNK_SIZE and CHUNK_OVERLAP from the worker’s environment if they are set, but they are not part of the deployment templates and changing them only affects documents processed afterwards. Existing chunks are not rebuilt.

3. Embedding

Each chunk is embedded with the embedding model bound to its collection or website:

  • Models: embedding models are registered per tenant in Admin > Models and enabled per space. There is no global default model in the environment.
  • Storage: vectors are stored in info_blob_chunks.embedding, a dimensionless pgvector column, so models with different dimensions coexist.
  • Provider calls: made through LiteLLM with the tenant’s configured provider credentials. See AI Provider Configuration.

4. Retrieval

When a user asks a question:

  1. The question is embedded with the same model as the assistant’s knowledge sources
  2. PostgreSQL orders chunks in scope (the assistant’s collections, websites, and integrations) by cosine distance and returns the nearest candidates: 30 in the legacy retrieval version, or a count derived from the model’s context window in the newer one
  3. Optionally, chunks below INJECT_KNOWLEDGE_MIN_SCORE are dropped
  4. In the legacy version an autocut step trims the low-relevance tail where the score curve flattens
  5. The remaining chunks are placed in the prompt as context

The score is 1 - cosine distance. Its scale depends on the embedding model, which is why INJECT_KNOWLEDGE_MIN_SCORE is unset by default.

See Knowledge Retrieval and MCP for the retrieval versions, the search-as-a-tool mode, and how the candidate count is derived from the model’s context window.


Optimizing Retrieval

Organise by Scope

Retrieval searches every chunk in the assistant’s knowledge scope, and there is no approximate-nearest-neighbour index on the embedding column. Smaller, focused collections are both faster and more precise than one large collection.

Choose the Embedding Model per Collection

The embedding model is fixed when a collection or website is created. To move content to another model, create a new collection with that model and upload the documents again; existing chunks are not re-embedded in place.

Relevance Floor

For inject-mode retrieval, set a deployment-wide floor in env_backend.env once you know the score distribution of your embedding model:

# Cosine similarity, -1 to 1. Unset by default. INJECT_KNOWLEDGE_MIN_SCORE=0.3

There are no other retrieval variables in the environment; the chunk count and autocut behaviour are set per retrieval version in code.


Managing Your Knowledge Base

View Documents

  1. Go to your space
  2. Click the Knowledge tab
  3. Open a collection or website to see its documents and job status

Update Documents

Upload the document again with the same title. Eneo builds the replacement first and makes it searchable only when its text and embeddings are complete. If processing fails, the previous version remains available.

Delete Documents

  1. Click the document
  2. Click Delete
  3. Confirm deletion

Deleting a document removes its text and all of its chunks and embeddings. Deleting a collection or website removes every document in it.


Performance

Worker Capacity

The worker runs the same image as the backend with RUN_AS_WORKER=true. Two settings in env_backend.env control throughput:

# Concurrent jobs per worker process (default 15) WORKER_MAX_JOBS=15 # Concurrent crawl jobs per tenant (default 4, 0 = unlimited) TENANT_WORKER_CONCURRENCY_LIMIT=4

Keep WORKER_MAX_JOBS at or below roughly 60 % of DB_POOL_SIZE + DB_POOL_MAX_OVERFLOW so the API keeps enough database connections. The backend template documents the sizing table. To add capacity, run an additional worker container with the same environment.

Health

  • GET /api/healthz reports the worker heartbeat; a stale heartbeat returns 503
  • GET /api/healthz/crawler shows crawl queue lengths and feeder diagnostics

Troubleshooting

Files Not Processing

Check worker status:

docker compose ps worker docker compose logs worker curl -s https://your-eneo-instance.com/api/healthz

Common issues:

  • Worker not running or its heartbeat is stale
  • File type not in the supported list, or a legacy .doc/.ppt file (unsupported_format)
  • Encrypted or corrupt file (the job shows encrypted or corrupt)
  • Scanned PDF with no text layer (no_extractable_text)
  • File larger than the limit in Admin > File storage

Solution:

docker compose restart worker

Web Crawling Fails

Check error messages:

docker compose logs worker | grep -i crawl curl -s https://your-eneo-instance.com/api/healthz/crawler

Common issues:

  • robots.txt disallows the crawler
  • The site requires HTTP Basic Authentication that was not configured
  • The crawl hit CLOSESPIDER_ITEMCOUNT or CRAWL_MAX_LENGTH
  • The website was auto-disabled after ten consecutive failures

Solution:

  • Use a sitemap crawl or a more specific start URL
  • Raise the tenant’s page or time limit through the sysadmin API
  • Change the website’s update interval to re-enable it

Poor Retrieval Quality

Symptoms:

  • AI doesn’t use uploaded documents
  • Irrelevant chunks retrieved
  • Missing important information

Solutions:

  1. Narrow the scope: attach only the collections the assistant needs
  2. Check the embedding model: all sources of one assistant must use a model that fits the content and language
  3. Set a relevance floor with INJECT_KNOWLEDGE_MIN_SCORE once you know your model’s score distribution
  4. Improve document quality: use text-based PDFs, remove boilerplate headers and footers, and keep one topic per document

Using the API

Every operation in this guide is available under /api/v1 and documented in the Swagger UI at /docs on your instance. Authenticate with an API key in the X-API-Key header.

Upload one file to a collection (one file per request; the response is a job you can poll):

import requests BASE = "https://your-eneo-instance.com/api/v1" HEADERS = {"X-API-Key": "your-api-key"} with open("doc1.pdf", "rb") as f: job = requests.post( f"{BASE}/groups/{collection_id}/info-blobs/upload/", files={"file": f}, headers=HEADERS, ).json() status = requests.get(f"{BASE}/jobs/{job['id']}/", headers=HEADERS).json()

Add text directly without a file: POST /api/v1/groups/{id}/info-blobs/ accepts up to 128 text info-blobs per request and embeds them with the collection’s model.

Create and run a website crawler:

website = requests.post( f"{BASE}/spaces/{space_id}/knowledge/websites/", json={ "url": "https://example.com/docs", "crawl_type": "crawl", # or "sitemap" "download_files": True, "update_interval": "weekly", # never | daily | every_other_day | weekly }, headers=HEADERS, ).json() run = requests.post(f"{BASE}/websites/{website['id']}/run/", headers=HEADERS).json()

See the API Reference for authentication details.


Best Practices

  1. Organize by topic: Create separate collections for different topics
  2. Use descriptive names: Name documents clearly; the title is what a re-upload matches on
  3. Keep documents updated: Re-upload changed files; the old version stays searchable until the new one is complete
  4. Monitor storage: Check disk usage periodically
  5. Test retrieval: Verify documents are being used in responses
  6. Start small: Test with a few documents before bulk upload
  7. Clean data: Remove unnecessary content before upload

Need Help?