Files
NSCT---Neutral-Search-Crawl…/ARCHITECTURE.md

288 lines
13 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# NSCT — Architekturdokumentation
## 1. Überblick
NSCT (Neutral Search Crawler Tool) ist ein modulares, containerisiertes
Recherchesystem. Es durchsucht Webquellen, extrahiert Inhalte, analysiert
Behauptungen (Claims) und erzeugt einen neutralen, quellengestützten Bericht.
**Kernprinzip:** Webcontent ist DATA, keine INSTRUCTION. Der LLM verarbeitet
nur strukturierte Daten, niemals rohen Webcontent direkt als Prompt.
**Stage 022 abgeschlossen.** NSCT ist production-ready.
---
## 2. Architektur-Diagramm
```
┌─────────────────────────────────────────────────────────────────────┐
│ Control Plane │
│ ┌──────────┐ ┌───────────┐ ┌────────────┐ ┌─────────────────┐ │
│ │ CLI │ │ REST API │ │ WebSocket │ │ Event Bus │ │
│ └────┬─────┘ └─────┬─────┘ └─────┬──────┘ └────────┬────────┘ │
│ │ │ │ │ │
│ └──────────────┴─────────────┴───────────────────┘ │
│ │
│ ┌───────────────────────Evidence Pipeline───────────────────────┐ │
│ │ │ │
│ │ ┌─────────┐ ┌──────────┐ ┌───────────┐ ┌───────────┐ │ │
│ │ │ Search │──▶│ Fetch │──▶│ Extract │──▶│ Classify │ │ │
│ │ │ (DDG/ │ │(HTTP/ │ │(Trafilat │ │(LLM/Rule)│ │ │
│ │ │ Multi) │ │ Playright)│ │ura/BS4) │ │ │ │ │
│ │ └─────────┘ └──────────┘ └───────────┘ └─────┬─────┘ │ │
│ │ │ │ │
│ │ ┌─────────┐ ┌───────────┐ ┌────────────┐ ┌───┴─────┐ │ │
│ │ │ Index │◀──│ Compare │◀──│ Normalise │◀──│ Claim │ │ │
│ │ │ & Store │ │ (LLM) │ │ & Parse │ │ Ext. │ │ │
│ │ └─────────┘ └─────┬─────┘ └────────────┘ └─────────┘ │ │
│ │ │ │ │
│ │ ▼ │ │
│ │ ┌──────────────┐ │ │
│ │ │Evidence │ │ │
│ │ │Scoring (6D) │ │ │
│ │ └──────┬───────┘ │ │
│ │ ▼ │ │
│ │ ┌──────────────┐ │ │
│ │ │ Report Gen. │ │ │
│ │ │ (LLM) │ │ │
│ │ └──────┬───────┘ │ │
│ │ ▼ │ │
│ │ ┌──────────────┐ │ │
│ │ │Provenance │ │ │
│ │ │+ Hash │ │ │
│ │ └──────────────┘ │ │
│ └──────────────────────────────────────────────────────────────┘ │
└─────────────────────────────────────────────────────────────────────┘
```
---
## 3. Datenfluss
```
Query ──▶ SearchProvider.search() ──▶ List[SearchResult]
│ (for each result)
ContentFetcher.fetch(url) ──▶ Source document (parsed)
│ (LLM: extract claims from source)
Claim[] ──▶ Normalization ──▶ NormalizedClaim[]
│ (LLM: compare claims pairwise)
EvidenceRelation[] ──▶ Agreement matrix
│ (6-dim evidence scoring)
EvidencePackage ──▶ Neutral Synthesis
│ (LLM: synthesize report)
ResearchReport (Summary + Findings + Disagreements + Uncertainties)
research_run_hash (deterministic, reproducible)
```
---
## 4. Module
### 4.1 `src/nsct/api/` — REST API
- **rest_research.py**: Research CRUD — POST/GET/DELETE + status, sources, claims, evidence, report
- **health.py**: `/health`, `/ready`, `/providers` endpunkte
- **search.py**: Search API endpoint
- **debug.py**: Debug-Endpunkte (nur bei `NSCT_DEBUG=true`)
- **Pattern:** Each router is a separate file; mounted in `main.py`
### 4.2 `src/nsct/models/` — Pydantic v2 Schemas
- **schemas.py**: Data transfer objects (SearchQuery, Source, Claim, etc.)
- **plan.py**: ResearchPlan Schema
- **claim.py**: Claim Schema mit Provenance
- **synthesis.py**: Synthese Report Schema
- **source_independence.py**: Independence Graph Schema
- **gap_analysis.py**: GapAnalysis Schema
- **audio.py**: Audio Transcription Schema
- **vision.py**: Vision Analysis Schema
- Pure Pydantic — no database or API coupling
### 4.3 `src/nsct/storage/` — Database Layer
- **models.py**: SQLAlchemy 2.0 declarative models
- **engine.py**: Async engine factory with pool management
- Uses `asyncpg` for PostgreSQL
### 4.4 `src/nsct/providers/` — Provider Interfaces
- **abstract.py**: Abstract base classes (LLMProvider, VisionProvider, SearchProvider, etc.)
- **llm.py**: OpenAI-kompatibler LLM Provider
- **duckduckgo.py**: DuckDuckGo Search Provider
- **multi.py**: MultiProviderSearch für parallele Abfrage
- **vision.py**: Qwen2.5-VL-3B Vision Provider
- **audio.py**: Audio/STT Provider
- **metrics.py**: Provider-Metriken Tracking
- **priority_queue.py**: Priority-Queue mit HIGH/NORMAL/LOW Prioritäten
- **semaphore.py**: Concurrency Semaphore (default: 3)
- **No concrete implementations in Stage 0** — just protocols
- All providers accept structured data, never raw web content
### 4.5 `src/nsct/security/` — Security
- **policy.py**: SSRF protection, URL validation, IP blocklisting
- Called before every outbound HTTP request
### 4.6 `src/nsct/config.py` — Configuration
- Pydantic BaseSettings — all values from environment
- Zero hard-coded secrets or URLs
- `AppSettings.from_env()` creates the root configuration
### 4.7 `src/nsct/logging_config.py` — Structured Logging
- JSON-formatted log output
- Per-request context tracking (`research_id`, `llm_request_id`)
- Global context dict merged into every log record
### 4.8 `src/nsct/metrics.py` — Prometheus Metrics
- Counters: search queries, sources fetched, claims, contradictions, completed/failed
- Histograms: research duration
- Gauges: active research runs
### 4.9 `src/nsct/agents/` — Research Planning
- **planner.py**: LLM-basierte Recherchestrategie (Query Expansion, Bias Reduction)
- **validator.py**: Plausibilitäts-Prüfung von Research-Plänen
### 4.10 `src/nsct/crawler/` — Content Extraction
- **fetcher.py**: Asynchroner HTTP-Fetcher mit SSRF-Schutz
- **extraction.py**: Main Content Extraction (trafilatura, BeautifulSoup)
- **normalize.py**: Dokumenten-Normalisierung
- **pdf.py**: PDF-Extraktion
- **policy.py**: SSRF/Download-Policy
- **manager.py**: Batch-Verwaltung
### 4.11 `src/nsct/orchestration/` — Pipeline Control
- **state.py**: State Machine mit 12 Zuständen
- **budget.py**: Hard Budget Limits (7 config options, frozen)
- **models.py**: ResearchRun Pydantic Model (frozen, immutable)
- **orchestrator.py**: Vollständige Pipeline-Steuerung mit Fallbacks
- **context_budget.py**: Pro-Stage Context Token Limits
- **gap_analysis.py**: Lückenerkennung für iterative Recherche
### 4.12 `src/nsct/stages/` — Pipeline Stages (5-13)
- **stage5_extract_claims.py**: Claim Extraction
- **stage6_source_independence.py**: Source Independence Graph
- **stage7_clustering.py**: Claim Clustering
- **stage7_normalize_numerics.py**: Numerical Normalization
- **stage8_evidence_scoring.py**: 6-dimensional Evidence Scoring
- **stage9_synthesis.py**: Neutral Synthesis
- **stage10_vision.py**: Vision Integration
- **stage11_audio.py**: Audio/STT Integration
- **stage13_gap_analysis.py**: Iterative Gap Analysis
### 4.13 `src/nsct/provenance.py` — Provenance & Reproducibility
- **Stage 21**: Vollständige Provenance aller Pipeline-Schritte
- `research_run_hash`: Deterministischer Hash für Reproduzierbarkeit
### 4.14 `src/nsct/cli.py` — Command Line Interface
- **Stage 15**: `nsct research`, `nsct status`, `nsct report`, etc.
---
## 5. Interfaces / Protocols
```python
# src/nsct/providers/__init__.py
class LLMProvider(ABC):
async def generate(prompt, system_prompt=None, ...) -> LLMResponse: ...
class VisionProvider(ABC):
async def analyze_image(image_bytes, prompt) -> VisionResponse: ...
class AudioProvider(ABC):
async def transcribe(audio_bytes) -> AudioResponse: ...
class SearchProvider(ABC):
async def search(query, language, limit) -> list[SearchResult]: ...
class ContentFetcher(ABC):
async def fetch(url, **kwargs) -> dict: ...
```
Alle konkreten Implementationen (DuckDuckGo, Qwen, Trafilatura, etc.)
müssen diese Interfaces implementieren — das ermöglicht den Wechsel
von Providern ohne Codeänderung im Core.
---
## 6. Sicherheitsarchitektur
- **Control Plane vs. Evidence Plane:** Der LLM verarbeitet nur strukturierte
Daten (Claims, NormalizedClaims, EvidenceRelations), niemals rohen Webcontent.
- **SSRF-Schutz:** Jede outbound-URL wird vor dem Request durch `validate_url()` geprüft.
- **Prompt-Isolation:** Webcontent wird niemals direkt als Prompt text eingebettet.
Stattdessen wird er in strukturierte JSON-Objekte serialisiert.
- **Non-Root-Docker:** Der Container läuft als nicht-root User `nsct`.
- **Read-Only-Filesystem:** Wo möglich (`read_only: true` + `tmpfs`).
- **Dropped Capabilities:** `cap_drop: [ALL]`
---
## 7. Datenmodell
### Core Tables
| Table | Beschreibung |
|-------|-------------|
| `search_queries` | Research-Aufträge mit Query, Purpose, Language |
| `sources` | Extrahierte Webquellen mit Metadaten |
| `claims` | Behauptungen aus Quellen mit Typ, Konfidenz, Evidence |
| `evidence_relations` | Vergleichsergebnisse zwischen Claims |
| `citation_edges` | Quelle-zu-Quelle Referenzen |
| `research_reports` | Aggregierte Forschungsberichte |
### Evidence Scoring (Stage 8)
6 dimensionale Scores für jede Evidenz:
| Dimension | Range | Beschreibung |
|-----------|-------|-------------|
| `source_independence` | 01 | Wie unabhängig ist diese Quelle? |
| `primary_source_proximity` | 01 | Wie nah ist die Quelle an der Primärquelle? |
| `cross_source_support` | 01 | Wie viele unabhängige Quellen bestätigen? |
| `contradiction_level` | 01 | Wie hoch ist der Widerspruch? (invertiert) |
| `evidence_directness` | 01 | Wie direkt ist die Evidenz? |
| `date_relevance_score` | 01 | Wie aktuell ist die Evidenz? |
---
## 8. Nicht-funktionale Anforderungen
| Kriterium | Anforderung |
|-----------|------------|
| Concurrency | Async I/O konsequent (async/await everywhere) |
| Database | PostgreSQL 16+, asyncpg, SQLAlchemy 2.0 |
| Logging | Strukturiert (JSON), jede Anfrage tracebar |
| Config | Environment-only, keine Config-Dateien |
| Testing | pytest-asyncio, FastAPI TestClient, E2E-Tests |
| Deployment | Docker Compose, reproduzierbar |
| LLM Concurrency | Konfigurierbar (`NSCT_LLM_MAX_CONCURRENCY`), default: 3 |
| Context Budgeting | Pro-Stage Limits (Planner→Claim→Contradiction→Synthesis) |
| Reproducibility | `research_run_hash` + vollständige Provenance |