Skip to content
AI·2026

RAG pipeline for internal knowledge bases

Ingestion from Confluence, SharePoint, and Markdown repos. Retrieval tuning with reranking, citation handling. Eval framework keeps hallucination below 1 %.

Methodology case6-Wochen-SprintpgvectorRAGCitations

Kontext / Context

Ein Beispiel-Setup für interne RAG-Implementierungen. Zeigt die wichtigsten Stellschrauben, die wir in realen Projekten konfigurieren, ohne spezifische Kundendetails.

Architektur

  • Ingestion: Connectors für Confluence, SharePoint, Markdown-Repos, PDF, DOCX, E-Mail. Inkrementelle Updates mit Change-Detection.
  • Chunking: Semantic-aware Chunking mit Overlap, Markdown-Heading-Anchored, Code-Block-Preserved.
  • Embeddings: voyage-3 oder text-embedding-3-large, mit Metadaten-Enrichment.
  • Storage: pgvector mit HNSW-Index, Postgres als Single-Source-of-Truth.
  • Retrieval: Hybrid (BM25 + Vector) mit Reranker (Cohere rerank-3).
  • Generation: claude-opus-4-7 für Antwort-Generierung, Citation-enforced Prompt.

Eval-Framework

  • Ground-Truth-Set: 200 Fragen mit Expert-Answers, inkrementell erweitert
  • Metriken: Faithfulness (grounded-in-context), Answer-Relevance, Context-Precision, Citation-Accuracy
  • Tool: Langfuse oder Braintrust für Traces und Regression-Gates
  • Daily-Run über Cron, Alert bei Regression auf Slack

Guardrails

  • PII-Detection im Query- und Answer-Layer
  • Confidence-Threshold: bei niedriger Similarity-Max “Ich weiß es nicht” statt Halluzination
  • Audit-Log pro Anfrage (Query, Retrieved, Answer, Citations, Latency)

Ready for a real assessment?

Short scoping call, clear proposal, kick-off in two to four weeks.

Start the conversation

Newsletter

Substance over noise

One email every two weeks with a new blog post or a technical deep dive. No clickbait.