hoangsonww/Customizable-AI-Chatbot

Knowledge Base Manager - In-App RAG Uploads, Versioning & Citations

Open

#12 opened on Aug 27, 2025

View on GitHub
 (0 comments) (0 reactions) (1 assignee)TypeScript (14 forks)auto 404
bugdocumentationenhancementgood first issuehelp wantedquestion

Repository metrics

Stars
 (22 stars)
PR merge metrics
 (PR metrics pending)

Description

🔎 Summary

Add a first-class, in-app Knowledge Base Manager so users can upload documents directly (PDF/MD/TXT/HTML), chunk → embed → upsert to Pinecone, manage versions, and get inline citations in chat. This removes the external dependency for RAG ingestion and makes the template truly “customizable” end-to-end.


🎯 Goals

  • Upload & manage sources in the UI; show status (queued, processing, indexed, error).
  • Deterministic chunking + metadata (title, url, checksum, version, page numbers).
  • Pluggable embedding provider (OpenAI by default; easy adapters for others).
  • Pinecone upsert in a per-project namespace; dedupe by checksum.
  • Chat answers include source snippets + links (with score threshold).
  • Safe, resumable processing with progress toasts + background jobs.
  • Export/import KB manifest (JSON) to migrate/backup.

🧱 Scope

In: Upload UI, server routes, chunk/embedding pipeline, Pinecone upsert, citations in chat, admin list & delete, namespace strategy, adapters, tests, docs. Out: Full OCR pipeline; we’ll parse PDFs (text layer) and skip image-only scans in v1 (graceful error / TODO: OCR).


🖥️ UX / UI

  • /kb page (protected or local-only toggle):

    • Dropzone (PDF/MD/TXT/HTML), multi-file upload.
    • Table: filename, size, type, checksum, version, status, chunks, updatedAt, actions (re-index, delete).
    • Filters: type, status, date.
    • Buttons: Add documents, Export manifest, Clear all (confirm).
  • Chat: When an answer includes sources, render a compact “Sources” footer with [title • p. X] badges; click opens side panel preview.


🧩 Data Model (Mongo or file-manifest JSON, pick your store)

kb_documents

  • _id, projectId (or single-tenant), filename, mime, size, checksum (SHA256), version (int), status (queued|processing|indexed|error), meta { title?, url?, pages? }, chunkCount, error?, createdAt, updatedAt.

kb_jobs

  • _id, documentId, type (ingest|reindex|delete), progress 0-100, log?, createdAt, updatedAt.

Pinecone

  • namespace = sha256(projectId || 'default')
  • id = ${checksum}:${chunkIndex}
  • metadata = { docId, filename, title, url, version, page, start, end }

🧠 Chunking & Embeddings

  • Chunking: token-aware splitter (approx 800-1200 tokens) with 200-token overlap; PDF includes page numbers.

  • Embeddings adapter interface:

    interface Embedder {
      model: string
      embed(texts: string[]): Promise<number[][]>
    }
    

    Default: OpenAI text-embedding-3-small. Adapter stubs: Fireworks, local (e5-small) via server endpoint (future).

  • Dedupe: If checksum seen with same version, skip. If file content changed → version++ and soft-delete old vectors (keep until reindex completes; then purge).


🔐 Security

  • Enforce size/type limits (env-configurable).
  • Scan/strip HTML; sanitize Markdown; ignore scripts.
  • Rate-limit uploads.
  • Do not log document content; redact file names if needed.

📈 Observability

  • Metrics: kb_files_total, kb_chunks_total, kb_index_errors_total.
  • Job progress events via SSE or polling (/api/kb/jobs/:id).

✅ Acceptance Criteria

  • Users can upload PDF/MD/TXT/HTML; see them indexed with chunk counts.
  • A chat that pulls context returns citations (>= top-k 3 with score ≤ threshold).
  • Re-uploading the same file (same bytes) does not duplicate vectors.
  • Updating a file re-indexes with version++ and purges prior vectors.
  • Deleting a doc removes its vectors from Pinecone within ≤ 2 minutes.
  • All flows covered by unit + integration tests.

🔌 API (Next.js Route Handlers)

  • POST /api/kb/upload — multipart; returns { documentId }.
  • GET /api/kb — list documents + statuses.
  • POST /api/kb/:id/reindex — reprocess a doc.
  • DELETE /api/kb/:id — delete doc + vectors.
  • GET /api/kb/jobs/:jobId — job status.
  • Chat retrieval: integrate in existing chat route: retrieve top-k by namespace, score filter, pass snippets to LLM, return citations.

⚙️ Config / Env

  • EMBEDDINGS_PROVIDER=openai
  • OPENAI_API_KEY=...
  • PINECONE_API_KEY=...
  • PINECONE_INDEX=my-ai
  • KB_MAX_UPLOAD_MB=25
  • KB_ALLOWED_MIME=application/pdf,text/markdown,text/plain,text/html
  • RAG_TOP_K=5
  • RAG_SCORE_THRESHOLD=0.35

🧪 Testing

  • Unit: chunk splitter, checksum, adapter interface, metadata mapping.
  • Integration: upload → index → chat w/ citations; update version; delete purge.
  • E2E (Playwright): happy path upload, progress UI, citations render.
  • Negative: oversize file, bad MIME, Pinecone error, partial failure resume.

📝 Tasks

  • UI: /kb page (dropzone, table, actions, progress toasts)
  • API: upload route (stream to temp, parse, checksum, enqueue job)
  • Worker: parse → chunk → embed → upsert (with backoff/retry)
  • API: list, reindex, delete, jobs status
  • Retrieval: augment chat route with top-k + score filter + citations payload
  • Pinecone: namespace + id conventions; purge old versions
  • Embedding adapter: OpenAI (default) + provider interface
  • Tests: unit/integration/E2E
  • Docs: README section “Knowledge Base Manager” (+ screenshots)
  • CI: size/type lint for fixtures, run tests on PR

🚀 Rollout

  1. Behind feature flag KB_MANAGER_ENABLED=1.
  2. Ship UI read-only list first (no upload) → then enable uploads.
  3. Monitor index errors; tune chunk size & thresholds.

Contributor guide