# NSCT — Architekturdokumentation ## 1. Überblick NSCT (Neutral Search Crawler Tool) ist ein modulares, containerisiertes Recherchesystem. Es durchsucht Webquellen, extrahiert Inhalte, analysiert Behauptungen (Claims) und erzeugt einen neutralen, quellengestützten Bericht. **Kernprinzip:** Webcontent ist DATA, keine INSTRUCTION. Der LLM verarbeitet nur strukturierte Daten, niemals rohen Webcontent direkt als Prompt. ## 2. Architektur-Diagramm ``` ┌─────────────────────────────────────────────────────────────────────┐ │ Control Plane │ │ ┌──────────┐ ┌───────────┐ ┌────────────┐ ┌─────────────────┐ │ │ │ CLI │ │ REST API │ │ WebSocket │ │ Event Bus │ │ │ └────┬─────┘ └─────┬─────┘ └─────┬──────┘ └────────┬────────┘ │ │ │ │ │ │ │ │ └──────────────┴─────────────┴───────────────────┘ │ │ │ │ ┌───────────────────────Evidence Pipeline───────────────────────┐ │ │ │ │ │ │ │ ┌─────────┐ ┌──────────┐ ┌───────────┐ ┌───────────┐ │ │ │ │ │ Search │──▶│ Fetch │──▶│ Extract │──▶│ Classify │ │ │ │ │ │ (SearX │ │ (HTTP/ │ │(Trafilat │ │(LLM/Rule)│ │ │ │ │ │ NG) │ │ Playwright)│ │ura/BS4) │ │ │ │ │ │ │ └─────────┘ └──────────┘ └───────────┘ └─────┬─────┘ │ │ │ │ │ │ │ │ │ ┌─────────┐ ┌───────────┐ ┌────────────┐ ┌───┴─────┐ │ │ │ │ │ Index │◀──│ Compare │◀──│ Normalise │◀──│ Claim │ │ │ │ │ │ & Store │ │ (LLM) │ │ & Parse │ │ Ext. │ │ │ │ │ └─────────┘ └─────┬─────┘ └────────────┘ └─────────┘ │ │ │ │ │ │ │ │ │ ▼ │ │ │ │ ┌──────────────┐ │ │ │ │ │ Report Gen. │ │ │ │ │ │ (LLM) │ │ │ │ │ └──────┬───────┘ │ │ │ │ ▼ │ │ │ │ ┌──────────────┐ │ │ │ │ │ Evidence DB │ │ │ │ │ └──────────────┘ │ │ │ └──────────────────────────────────────────────────────────────┘ │ └─────────────────────────────────────────────────────────────────────┘ ``` ## 3. Datenfluss ``` Query ──▶ SearchProvider.search() ──▶ List[SearchResult] │ │ (for each result) │ ▼ ContentFetcher.fetch(url) ──▶ Source document (parsed) │ │ (LLM: extract claims from source) │ ▼ Claim[] ──▶ Normalization ──▶ NormalizedClaim[] │ │ (LLM: compare claims pairwise) │ ▼ EvidenceRelation[] ──▶ Agreement matrix │ │ (LLM: synthesize report) │ ▼ ResearchReport (Summary + Findings + Disagreements + Uncertainties) ``` ## 4. Module ### 4.1 `src/nsct/api/` — REST API - **main.py**: FastAPI application factory, CORS, lifespan hooks - **health.py**: `/health`, `/ready`, `/providers` endpunkte - **Pattern:** Each router is a separate file; mounted in `main.py` ### 4.2 `src/nsct/models/` — Pydantic v2 Schemas - **schemas.py**: Data transfer objects (SearchQuery, Source, Claim, etc.) - Pure Pydantic — no database or API coupling ### 4.3 `src/nsct/storage/` — Database Layer - **models.py**: SQLAlchemy 2.0 declarative models - **engine.py**: Async engine factory with pool management - Uses `asyncpg` for PostgreSQL ### 4.4 `src/nsct/providers/` — Provider Interfaces - Abstract base classes (LLMProvider, VisionProvider, SearchProvider, ContentFetcher, AudioProvider) - **No concrete implementations in Stage 0** — just protocols - All providers accept structured data, never raw web content ### 4.5 `src/nsct/security/` — Security - **policy.py**: SSRF protection, URL validation, IP blocklisting - Called before every outbound HTTP request ### 4.6 `src/nsct/config.py` — Configuration - Pydantic BaseSettings — all values from environment - Zero hard-coded secrets or URLs - `AppSettings.from_env()` creates the root configuration ### 4.7 `src/nsct/logging_config.py` — Structured Logging - JSON-formatted log output - Per-request context tracking (`research_id`, `llm_request_id`) - Global context dict merged into every log record ## 5. Interfaces / Protocols ```python # src/nsct/providers/__init__.py class LLMProvider(ABC): async def generate(prompt, system_prompt=None, ...) -> LLMResponse: ... class VisionProvider(ABC): async def analyze_image(image_bytes, prompt) -> VisionResponse: ... class AudioProvider(ABC): async def transcribe(audio_bytes) -> AudioResponse: ... class SearchProvider(ABC): async def search(query, language, limit) -> list[SearchResult]: ... class ContentFetcher(ABC): async def fetch(url, **kwargs) -> dict: ... ``` Alle konkreten Implementationen (SearXNG, OpenAI, Trafilatura, etc.) müssen diese Interfaces implementieren — das ermöglicht den Wechsel von Providern ohne Codeänderung im Core. ## 6. Sicherheitsarchitektur - **Control Plane vs. Evidence Plane:** Der LLM verarbeitet nur strukturierte Daten (Claims, NormalizedClaims, EvidenceRelations), niemals rohen Webcontent. - **SSRF-Schutz:** Jede outbound-URL wird vor dem Request durch `validate_url()` geprüft. - **Prompt-Isolation:** Webcontent wird niemals direkt als Prompt text eingebettet. Stattdessen wird er in strukturierte JSON-Objekte serialisiert. - **Non-Root-Docker:** Der Container läuft als nicht-root User `nsct`. - **Read-Only-Filesystem:** Wo möglich (`read_only: true` + `tmpfs`). ## 7. Datenmodell ### Core Tables | Table | Beschreibung | |-------|-------------| | `search_queries` | Research-Aufträge mit Query, Purpose, Language | | `sources` | Extrahierte Webquellen mit Metadaten | | `claims` | Behauptungen aus Quellen mit Typ, Konfidenz, Evidence | | `evidence_relations` | Vergleichsergebnisse zwischen Claims | | `citation_edges` | Quelle-zu-Quelle Referenzen | | `research_reports` | Aggregierte Forschungsberichte | ## 8. Nicht-funktionale Anforderungen | Kriterium | Anforderung | |-----------|------------| | Concurrency | Async I/O konsequent (async/await everywhere) | | Database | PostgreSQL 16+, asyncpg, SQLAlchemy 2.0 | | Logging | Strukturiert (JSON), jede Anfrage tracebar | | Config | Environment-only, keine Config-Dateien | | Testing | pytest-asyncio, FastAPI TestClient | | Deployment | Docker Compose, reproduzierbar |