288 lines
13 KiB
Markdown
288 lines
13 KiB
Markdown
# NSCT — Architekturdokumentation
|
||
|
||
## 1. Überblick
|
||
|
||
NSCT (Neutral Search Crawler Tool) ist ein modulares, containerisiertes
|
||
Recherchesystem. Es durchsucht Webquellen, extrahiert Inhalte, analysiert
|
||
Behauptungen (Claims) und erzeugt einen neutralen, quellengestützten Bericht.
|
||
|
||
**Kernprinzip:** Webcontent ist DATA, keine INSTRUCTION. Der LLM verarbeitet
|
||
nur strukturierte Daten, niemals rohen Webcontent direkt als Prompt.
|
||
|
||
**Stage 0–22 abgeschlossen.** NSCT ist production-ready.
|
||
|
||
---
|
||
|
||
## 2. Architektur-Diagramm
|
||
|
||
```
|
||
┌─────────────────────────────────────────────────────────────────────┐
|
||
│ Control Plane │
|
||
│ ┌──────────┐ ┌───────────┐ ┌────────────┐ ┌─────────────────┐ │
|
||
│ │ CLI │ │ REST API │ │ WebSocket │ │ Event Bus │ │
|
||
│ └────┬─────┘ └─────┬─────┘ └─────┬──────┘ └────────┬────────┘ │
|
||
│ │ │ │ │ │
|
||
│ └──────────────┴─────────────┴───────────────────┘ │
|
||
│ │
|
||
│ ┌───────────────────────Evidence Pipeline───────────────────────┐ │
|
||
│ │ │ │
|
||
│ │ ┌─────────┐ ┌──────────┐ ┌───────────┐ ┌───────────┐ │ │
|
||
│ │ │ Search │──▶│ Fetch │──▶│ Extract │──▶│ Classify │ │ │
|
||
│ │ │ (DDG/ │ │(HTTP/ │ │(Trafilat │ │(LLM/Rule)│ │ │
|
||
│ │ │ Multi) │ │ Playright)│ │ura/BS4) │ │ │ │ │
|
||
│ │ └─────────┘ └──────────┘ └───────────┘ └─────┬─────┘ │ │
|
||
│ │ │ │ │
|
||
│ │ ┌─────────┐ ┌───────────┐ ┌────────────┐ ┌───┴─────┐ │ │
|
||
│ │ │ Index │◀──│ Compare │◀──│ Normalise │◀──│ Claim │ │ │
|
||
│ │ │ & Store │ │ (LLM) │ │ & Parse │ │ Ext. │ │ │
|
||
│ │ └─────────┘ └─────┬─────┘ └────────────┘ └─────────┘ │ │
|
||
│ │ │ │ │
|
||
│ │ ▼ │ │
|
||
│ │ ┌──────────────┐ │ │
|
||
│ │ │Evidence │ │ │
|
||
│ │ │Scoring (6D) │ │ │
|
||
│ │ └──────┬───────┘ │ │
|
||
│ │ ▼ │ │
|
||
│ │ ┌──────────────┐ │ │
|
||
│ │ │ Report Gen. │ │ │
|
||
│ │ │ (LLM) │ │ │
|
||
│ │ └──────┬───────┘ │ │
|
||
│ │ ▼ │ │
|
||
│ │ ┌──────────────┐ │ │
|
||
│ │ │Provenance │ │ │
|
||
│ │ │+ Hash │ │ │
|
||
│ │ └──────────────┘ │ │
|
||
│ └──────────────────────────────────────────────────────────────┘ │
|
||
└─────────────────────────────────────────────────────────────────────┘
|
||
```
|
||
|
||
---
|
||
|
||
## 3. Datenfluss
|
||
|
||
```
|
||
Query ──▶ SearchProvider.search() ──▶ List[SearchResult]
|
||
│
|
||
│ (for each result)
|
||
│
|
||
▼
|
||
ContentFetcher.fetch(url) ──▶ Source document (parsed)
|
||
│
|
||
│ (LLM: extract claims from source)
|
||
│
|
||
▼
|
||
Claim[] ──▶ Normalization ──▶ NormalizedClaim[]
|
||
│
|
||
│ (LLM: compare claims pairwise)
|
||
│
|
||
▼
|
||
EvidenceRelation[] ──▶ Agreement matrix
|
||
│
|
||
│ (6-dim evidence scoring)
|
||
│
|
||
▼
|
||
EvidencePackage ──▶ Neutral Synthesis
|
||
│
|
||
│ (LLM: synthesize report)
|
||
│
|
||
▼
|
||
ResearchReport (Summary + Findings + Disagreements + Uncertainties)
|
||
│
|
||
▼
|
||
research_run_hash (deterministic, reproducible)
|
||
```
|
||
|
||
---
|
||
|
||
## 4. Module
|
||
|
||
### 4.1 `src/nsct/api/` — REST API
|
||
|
||
- **rest_research.py**: Research CRUD — POST/GET/DELETE + status, sources, claims, evidence, report
|
||
- **health.py**: `/health`, `/ready`, `/providers` endpunkte
|
||
- **search.py**: Search API endpoint
|
||
- **debug.py**: Debug-Endpunkte (nur bei `NSCT_DEBUG=true`)
|
||
- **Pattern:** Each router is a separate file; mounted in `main.py`
|
||
|
||
### 4.2 `src/nsct/models/` — Pydantic v2 Schemas
|
||
|
||
- **schemas.py**: Data transfer objects (SearchQuery, Source, Claim, etc.)
|
||
- **plan.py**: ResearchPlan Schema
|
||
- **claim.py**: Claim Schema mit Provenance
|
||
- **synthesis.py**: Synthese Report Schema
|
||
- **source_independence.py**: Independence Graph Schema
|
||
- **gap_analysis.py**: GapAnalysis Schema
|
||
- **audio.py**: Audio Transcription Schema
|
||
- **vision.py**: Vision Analysis Schema
|
||
- Pure Pydantic — no database or API coupling
|
||
|
||
### 4.3 `src/nsct/storage/` — Database Layer
|
||
|
||
- **models.py**: SQLAlchemy 2.0 declarative models
|
||
- **engine.py**: Async engine factory with pool management
|
||
- Uses `asyncpg` for PostgreSQL
|
||
|
||
### 4.4 `src/nsct/providers/` — Provider Interfaces
|
||
|
||
- **abstract.py**: Abstract base classes (LLMProvider, VisionProvider, SearchProvider, etc.)
|
||
- **llm.py**: OpenAI-kompatibler LLM Provider
|
||
- **duckduckgo.py**: DuckDuckGo Search Provider
|
||
- **multi.py**: MultiProviderSearch für parallele Abfrage
|
||
- **vision.py**: Qwen2.5-VL-3B Vision Provider
|
||
- **audio.py**: Audio/STT Provider
|
||
- **metrics.py**: Provider-Metriken Tracking
|
||
- **priority_queue.py**: Priority-Queue mit HIGH/NORMAL/LOW Prioritäten
|
||
- **semaphore.py**: Concurrency Semaphore (default: 3)
|
||
- **No concrete implementations in Stage 0** — just protocols
|
||
- All providers accept structured data, never raw web content
|
||
|
||
### 4.5 `src/nsct/security/` — Security
|
||
|
||
- **policy.py**: SSRF protection, URL validation, IP blocklisting
|
||
- Called before every outbound HTTP request
|
||
|
||
### 4.6 `src/nsct/config.py` — Configuration
|
||
|
||
- Pydantic BaseSettings — all values from environment
|
||
- Zero hard-coded secrets or URLs
|
||
- `AppSettings.from_env()` creates the root configuration
|
||
|
||
### 4.7 `src/nsct/logging_config.py` — Structured Logging
|
||
|
||
- JSON-formatted log output
|
||
- Per-request context tracking (`research_id`, `llm_request_id`)
|
||
- Global context dict merged into every log record
|
||
|
||
### 4.8 `src/nsct/metrics.py` — Prometheus Metrics
|
||
|
||
- Counters: search queries, sources fetched, claims, contradictions, completed/failed
|
||
- Histograms: research duration
|
||
- Gauges: active research runs
|
||
|
||
### 4.9 `src/nsct/agents/` — Research Planning
|
||
|
||
- **planner.py**: LLM-basierte Recherchestrategie (Query Expansion, Bias Reduction)
|
||
- **validator.py**: Plausibilitäts-Prüfung von Research-Plänen
|
||
|
||
### 4.10 `src/nsct/crawler/` — Content Extraction
|
||
|
||
- **fetcher.py**: Asynchroner HTTP-Fetcher mit SSRF-Schutz
|
||
- **extraction.py**: Main Content Extraction (trafilatura, BeautifulSoup)
|
||
- **normalize.py**: Dokumenten-Normalisierung
|
||
- **pdf.py**: PDF-Extraktion
|
||
- **policy.py**: SSRF/Download-Policy
|
||
- **manager.py**: Batch-Verwaltung
|
||
|
||
### 4.11 `src/nsct/orchestration/` — Pipeline Control
|
||
|
||
- **state.py**: State Machine mit 12 Zuständen
|
||
- **budget.py**: Hard Budget Limits (7 config options, frozen)
|
||
- **models.py**: ResearchRun Pydantic Model (frozen, immutable)
|
||
- **orchestrator.py**: Vollständige Pipeline-Steuerung mit Fallbacks
|
||
- **context_budget.py**: Pro-Stage Context Token Limits
|
||
- **gap_analysis.py**: Lückenerkennung für iterative Recherche
|
||
|
||
### 4.12 `src/nsct/stages/` — Pipeline Stages (5-13)
|
||
|
||
- **stage5_extract_claims.py**: Claim Extraction
|
||
- **stage6_source_independence.py**: Source Independence Graph
|
||
- **stage7_clustering.py**: Claim Clustering
|
||
- **stage7_normalize_numerics.py**: Numerical Normalization
|
||
- **stage8_evidence_scoring.py**: 6-dimensional Evidence Scoring
|
||
- **stage9_synthesis.py**: Neutral Synthesis
|
||
- **stage10_vision.py**: Vision Integration
|
||
- **stage11_audio.py**: Audio/STT Integration
|
||
- **stage13_gap_analysis.py**: Iterative Gap Analysis
|
||
|
||
### 4.13 `src/nsct/provenance.py` — Provenance & Reproducibility
|
||
|
||
- **Stage 21**: Vollständige Provenance aller Pipeline-Schritte
|
||
- `research_run_hash`: Deterministischer Hash für Reproduzierbarkeit
|
||
|
||
### 4.14 `src/nsct/cli.py` — Command Line Interface
|
||
|
||
- **Stage 15**: `nsct research`, `nsct status`, `nsct report`, etc.
|
||
|
||
---
|
||
|
||
## 5. Interfaces / Protocols
|
||
|
||
```python
|
||
# src/nsct/providers/__init__.py
|
||
|
||
class LLMProvider(ABC):
|
||
async def generate(prompt, system_prompt=None, ...) -> LLMResponse: ...
|
||
|
||
class VisionProvider(ABC):
|
||
async def analyze_image(image_bytes, prompt) -> VisionResponse: ...
|
||
|
||
class AudioProvider(ABC):
|
||
async def transcribe(audio_bytes) -> AudioResponse: ...
|
||
|
||
class SearchProvider(ABC):
|
||
async def search(query, language, limit) -> list[SearchResult]: ...
|
||
|
||
class ContentFetcher(ABC):
|
||
async def fetch(url, **kwargs) -> dict: ...
|
||
```
|
||
|
||
Alle konkreten Implementationen (DuckDuckGo, Qwen, Trafilatura, etc.)
|
||
müssen diese Interfaces implementieren — das ermöglicht den Wechsel
|
||
von Providern ohne Codeänderung im Core.
|
||
|
||
---
|
||
|
||
## 6. Sicherheitsarchitektur
|
||
|
||
- **Control Plane vs. Evidence Plane:** Der LLM verarbeitet nur strukturierte
|
||
Daten (Claims, NormalizedClaims, EvidenceRelations), niemals rohen Webcontent.
|
||
- **SSRF-Schutz:** Jede outbound-URL wird vor dem Request durch `validate_url()` geprüft.
|
||
- **Prompt-Isolation:** Webcontent wird niemals direkt als Prompt text eingebettet.
|
||
Stattdessen wird er in strukturierte JSON-Objekte serialisiert.
|
||
- **Non-Root-Docker:** Der Container läuft als nicht-root User `nsct`.
|
||
- **Read-Only-Filesystem:** Wo möglich (`read_only: true` + `tmpfs`).
|
||
- **Dropped Capabilities:** `cap_drop: [ALL]`
|
||
|
||
---
|
||
|
||
## 7. Datenmodell
|
||
|
||
### Core Tables
|
||
|
||
| Table | Beschreibung |
|
||
|-------|-------------|
|
||
| `search_queries` | Research-Aufträge mit Query, Purpose, Language |
|
||
| `sources` | Extrahierte Webquellen mit Metadaten |
|
||
| `claims` | Behauptungen aus Quellen mit Typ, Konfidenz, Evidence |
|
||
| `evidence_relations` | Vergleichsergebnisse zwischen Claims |
|
||
| `citation_edges` | Quelle-zu-Quelle Referenzen |
|
||
| `research_reports` | Aggregierte Forschungsberichte |
|
||
|
||
### Evidence Scoring (Stage 8)
|
||
|
||
6 dimensionale Scores für jede Evidenz:
|
||
|
||
| Dimension | Range | Beschreibung |
|
||
|-----------|-------|-------------|
|
||
| `source_independence` | 0–1 | Wie unabhängig ist diese Quelle? |
|
||
| `primary_source_proximity` | 0–1 | Wie nah ist die Quelle an der Primärquelle? |
|
||
| `cross_source_support` | 0–1 | Wie viele unabhängige Quellen bestätigen? |
|
||
| `contradiction_level` | 0–1 | Wie hoch ist der Widerspruch? (invertiert) |
|
||
| `evidence_directness` | 0–1 | Wie direkt ist die Evidenz? |
|
||
| `date_relevance_score` | 0–1 | Wie aktuell ist die Evidenz? |
|
||
|
||
---
|
||
|
||
## 8. Nicht-funktionale Anforderungen
|
||
|
||
| Kriterium | Anforderung |
|
||
|-----------|------------|
|
||
| Concurrency | Async I/O konsequent (async/await everywhere) |
|
||
| Database | PostgreSQL 16+, asyncpg, SQLAlchemy 2.0 |
|
||
| Logging | Strukturiert (JSON), jede Anfrage tracebar |
|
||
| Config | Environment-only, keine Config-Dateien |
|
||
| Testing | pytest-asyncio, FastAPI TestClient, E2E-Tests |
|
||
| Deployment | Docker Compose, reproduzierbar |
|
||
| LLM Concurrency | Konfigurierbar (`NSCT_LLM_MAX_CONCURRENCY`), default: 3 |
|
||
| Context Budgeting | Pro-Stage Limits (Planner→Claim→Contradiction→Synthesis) |
|
||
| Reproducibility | `research_run_hash` + vollständige Provenance | |