Files
curo1305andClaude Sonnet 4.6 3df62506c9 docs(07): create phase plan — 5 plans, verification passed
5 waves: system_settings DB + HKDF encryption (01), ProviderConfig +
GenericOpenAIProvider + singleton client fix (02), Anthropic output_config +
classifier wiring (03), Celery retry 30/90/270s + re-queue endpoint (04),
admin AI panel + DocumentCard badge + human checkpoint (05).

All 18 decisions D-01..D-18 covered. Plan checker passed after 1 revision
round (4 blockers fixed: D-02/D-14 coverage, D-11 Vitest tests, Plan 05
files_modified).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-03 18:25:00 +02:00

55 KiB

Phase 7: Redo and Optimize LLM Integration - Research

Researched: 2026-06-02 Domain: AI provider abstraction layer, Celery retry, structured output, DB-driven config Confidence: HIGH


<user_constraints>

User Constraints (from CONTEXT.md)

Locked Decisions

  • D-01: All OpenAI-compatible providers (OpenAI, Ollama, LMStudio, Groq, Gemini-compat, Mistral-compat, xAI/Grok, DeepSeek, OpenRouter) must pass response_format={"type": "json_object"} in every classify() and suggest_topics() call.
  • D-02: Keep parse_classification() and parse_suggestions() in ai/utils.py as a last-resort fallback only — they are called when a provider does not support response_format (e.g., some local Ollama models that ignore the parameter).
  • D-04: All AI provider settings (API keys, base URLs, model names, active provider) are stored in a new system_settings DB table. Env vars serve as startup defaults that populate the table on first boot, but the DB record is authoritative at runtime.
  • D-05: API keys stored in system_settings must be encrypted at rest — use the same HKDF/AES-GCM pattern already established for cloud credentials in backend/storage/cloud_utils.py.
  • D-06: get_provider() accepts typed Pydantic ProviderConfig models (not raw dicts). The classifier assembles a ProviderConfig from the DB record.
  • D-08: Admin panel gains an "AI Providers" configuration section where admins can set the active provider, configure API keys/base URLs/models, and test connectivity via the existing health_check() method.
  • D-09: The extract_and_classify Celery task uses Celery's built-in retry mechanism (self.retry(exc=…, countdown=…, max_retries=3)) with exponential backoff: first retry at 30 s, second at 90 s, third at 270 s.
  • D-10: After all retries are exhausted the document status remains classification_failed. No silent failure.
  • D-11: The document card shows a "Classification failed" badge when status == "classification_failed". A "Re-analyze" button calls POST /api/documents/{id}/classify to re-queue the Celery task.
  • D-12: MAX_AI_CHARS is replaced by a per-provider context_chars value declared in each provider's config (or ProviderConfig default). The global constant is removed.
  • D-13: Smart truncation strategy: take the first 60% of the per-provider limit + the last 40%.
  • D-14: Add extra_hosts: ["host.docker.internal:host-gateway"] to the backend and celery-worker services in docker-compose.yml.
  • D-15: LMStudio base URL and model are configurable via the admin AI Providers panel.
  • D-16: Introduce a GenericOpenAIProvider(base_url, api_key, model, context_chars) class in backend/ai/generic_openai_provider.py. It subclasses OpenAIProvider and adds JSON-mode enforcement.
  • D-17: Providers to add: Google Gemini (via OpenAI-compat endpoint), Mistral (via OpenAI-compat), Groq, xAI/Grok, DeepSeek, OpenRouter — all via GenericOpenAIProvider.
  • D-18: Mistral uses the OpenAI-compat endpoint (not the native mistralai SDK).

Claude's Discretion

  • D-03 (Anthropic output): Researcher determines the most reliable pattern for structured JSON from Claude models consistent with D-01.
  • D-07 (Client lifecycle): Researcher documents optimal client instantiation per provider (singleton, per-request, pooled) clearly enough that adding a new provider requires no re-research.

Deferred Ideas (OUT OF SCOPE)

  • Chunk + merge multi-call classification (too expensive per-API call in this phase)
  • Native mistralai SDK (OpenAI-compat is sufficient)
  • Token counting with tiktoken or provider tokenizers
  • Streaming classification

</user_constraints>


Summary

Phase 7 is a focused refactor of an already-functional AI layer. The goal is reliability and extensibility: move provider settings from env vars into a DB table (system_settings), introduce a GenericOpenAIProvider that covers all OpenAI-compatible APIs in one class, enforce JSON-mode structured output across all providers, add Celery retry with exponential backoff for classification failures, and surface a "Re-analyze" button in the UI when classification fails.

The research resolves the two open design questions: (1) Anthropic now supports native output_config.format.type="json_schema" on all current Claude models with constrained decoding — no tool_use workaround is needed; (2) the optimal client lifecycle pattern is a per-provider module-level singleton that is constructed once at provider instantiation and stored as self._client, not recreated per API call.

A critical codebase finding: extra_hosts: ["host.docker.internal:host-gateway"] is already present on both backend and celery-worker in docker-compose.yml — D-14 is already done. Similarly, POST /api/documents/{id}/classify already exists in api/documents.py and classifyDocument() is already in the frontend API client. The phase needs to change what happens inside that endpoint (re-queue Celery instead of running synchronously) and update the frontend badge logic, not add new routing from scratch.

Primary recommendation: Implement in five waves: (1) Alembic migration + system_settings model + encryption helpers, (2) ProviderConfig Pydantic models + GenericOpenAIProvider + refactored provider constructors, (3) Anthropic output_config structured output + classifier refactor with smart truncation, (4) Celery retry harness + re-classify endpoint to re-queue, (5) Admin AI Config panel UI + DocumentCard badge + tests.


Architectural Responsibility Map

Capability Primary Tier Secondary Tier Rationale
system_settings read/write Database (PostgreSQL ORM) API / Backend Settings are global admin config, live in DB
API key encryption/decryption API / Backend (services layer) Crypto must never happen in the DB or frontend
Provider factory / get_provider() API / Backend (ai/__init__.py) Instantiation is a backend concern
JSON structured output enforcement API / Backend (each provider class) Each provider's classify() sets response_format
Context window truncation API / Backend (each provider class) Done before calling the provider API
Celery retry logic Celery worker (tasks/document_tasks.py) Task retry is a worker-level concern
Re-classify trigger API / Backend endpoint Frontend (button) Backend re-queues; frontend emits the call
Admin AI config UI Frontend (AdminAiConfigTab.vue) Admin API endpoint UI displays/saves; backend enforces
Classification failed badge Frontend (DocumentCard.vue) UI reads doc.status == "classification_failed"

Standard Stack

Core (no new packages — all already in requirements.txt)

Library Version (pinned min) Purpose Why Standard
openai >=1.30 (latest: 2.40.0) AsyncOpenAI for OpenAI + all compat providers Official SDK; manages httpx connection pool
anthropic >=0.26 (latest: 0.105.2) AsyncAnthropic for native Claude output Official SDK; supports output_config.format
celery[redis] >=5.5.0 (latest: 5.6.3) Task retry with exponential backoff Already used for document processing
cryptography >=41.0.0 HKDF/AES-GCM for system_settings API key encryption Same pattern as cloud credentials
pydantic >=2.0 ProviderConfig typed models Already used everywhere
sqlalchemy[asyncio] >=2.0.49 SystemSettings ORM model Already used everywhere
httpx >=0.27 (latest: 0.28.1) Transport layer for all async HTTP clients Used internally by both openai and anthropic SDKs

No new packages need to be added to requirements.txt for this phase. All required libraries are already present.

Alternatives Considered

Instead of Could Use Tradeoff
openai SDK for compat providers httpx directly openai SDK manages connection pooling, retries, and auth headers automatically — less code
output_config.format for Anthropic tool_use JSON schema output_config is simpler, no tool_call parsing needed, and uses constrained decoding
Per-provider context_chars int tiktoken token counts Char approximation is 4x cheaper in CPU and requires no extra dependency; deferred in CONTEXT.md

Package Legitimacy Audit

No new packages are introduced in this phase. All libraries (openai, anthropic, celery, cryptography, pydantic, sqlalchemy, httpx) are long-established, present in requirements.txt, and well-known to the ecosystem.

Package Registry Age Status
openai PyPI ~5 yrs Established — official OpenAI Python SDK
anthropic PyPI ~3 yrs Established — official Anthropic Python SDK
celery PyPI ~15 yrs Established
cryptography PyPI ~10 yrs Established

Packages removed due to slopcheck verdict: none — no new packages.


Architecture Patterns

System Architecture Diagram

[Admin Browser]
    │ PUT /api/admin/ai-config
    ▼
[Admin API (admin.py)]
    │ HKDF-encrypt(api_key)  write to system_settings table
    ▼
[PostgreSQL: system_settings table]
    │ load_provider_config(session) reads on each classify call
    ▼
[classifier.py — classify_document()]
    │ builds ProviderConfig  calls get_provider(config)
    ▼
[ai/__init__.py — get_provider(config: ProviderConfig)]
    │ registry lookup  construct AnthropicProvider | GenericOpenAIProvider
    ▼
[Provider instance (singleton _client)]
    │ classify() / suggest_topics()  send request to LLM API
    ▼
[External LLM API]
    │ response JSON (or fallback parse_classification())
    ▼
[classifier.py] — persist topics to DB
    ▲
[Celery task: extract_and_classify (bind=True)]
    │ on exception: self.retry(countdown=30|90|270, max_retries=3)
    │ on exhaustion: doc.status = "classification_failed"
    ▼
[Document: status field]
    │ polled by frontend
    ▼
[DocumentCard.vue]
    └── classification_failed badge + Re-analyze button
              │ POST /api/documents/{id}/classify
              ▼
        [documents.py — re-queues Celery task]
backend/
├── ai/
│   ├── base.py                        # AIProvider ABC (no change)
│   ├── utils.py                       # parse_classification/parse_suggestions (no change)
│   ├── __init__.py                    # get_provider(config: ProviderConfig) — refactored
│   ├── provider_config.py             # NEW: ProviderConfig Pydantic model + presets
│   ├── generic_openai_provider.py     # NEW: GenericOpenAIProvider(OpenAIProvider)
│   ├── openai_provider.py             # ADD context_chars, singleton _client, json_mode
│   ├── anthropic_provider.py          # ADD output_config structured output, singleton _client
│   ├── ollama_provider.py             # UPDATE: pass context_chars
│   └── lmstudio_provider.py           # UPDATE: pass context_chars
├── db/
│   └── models.py                      # ADD SystemSettings ORM model
├── alembic/versions/
│   └── 0005_system_settings.py        # NEW: migration for system_settings table
├── services/
│   └── ai_config.py                   # NEW: load_provider_config(session) helper
└── tasks/
    └── document_tasks.py              # REFACTOR: bind=True, self.retry() exponential backoff

frontend/src/
├── components/
│   ├── admin/
│   │   └── AdminAiConfigTab.vue       # REPLACE: global provider config (not per-user)
│   └── documents/
│       └── DocumentCard.vue           # ADD: classification_failed badge + Re-analyze button
└── api/
    └── client.js                      # ADD: reclassifyDocument(), getAiConfig(), saveAiConfig()

Don't Hand-Roll

Problem Don't Build Use Instead Why
JSON schema enforcement on LLM output Custom regex/JSON repair response_format={"type":"json_object"} (OpenAI compat) or output_config.format (Anthropic) SDK-level constraint; cannot produce invalid JSON with constrained decoding
Exponential backoff with jitter Manual sleep loop self.retry(countdown=N, max_retries=3) in Celery Celery handles scheduling, worker redelivery, and failure state atomically
Per-provider connection pools Custom httpx session management Module-level AsyncOpenAI / AsyncAnthropic singletons Both SDKs manage httpx connection pools internally; re-instantiation destroys pool reuse
HKDF key derivation AES with hardcoded key _derive_fernet_key(master_key, provider_id) from storage/cloud_utils.py Pattern already battle-tested in Phase 5; reuse verbatim
Provider capability detection Runtime feature probe Static supports_json_mode: bool flag per preset Avoids extra API calls; Ollama/LMStudio's fallback is already handled in D-02

Key insight: The entire OpenAI-compat ecosystem (Groq, xAI, DeepSeek, OpenRouter, Gemini-compat, Mistral-compat) is handled by a single GenericOpenAIProvider with different base_url defaults. The JSON-mode parameter travels through unchanged. Do not implement separate classes for each.


D-03 Resolved: Anthropic Structured Output Strategy

Decision: Use output_config={"format": {"type": "json_schema", "schema": {...}}} in messages.create().

Rationale:

  • The Anthropic Python SDK (v0.95+) supports output_config.format.type = "json_schema" with constrained decoding — the model is physically prevented from producing invalid JSON matching the schema. [VERIFIED: platform.claude.com/docs/en/build-with-claude/structured-outputs]
  • Supported on all current Claude models: Claude Haiku 4.5, Sonnet 4.5, Sonnet 4.6, Opus 4.5, Opus 4.6, Opus 4.7, Opus 4.8. [VERIFIED: platform.claude.com/docs/en/build-with-claude/structured-outputs]
  • No beta headers are required as of SDK ≥0.95 (the old structured-outputs-2025-11-13 header is still accepted for backward compat but no longer needed). [VERIFIED: platform.claude.com/docs/en/build-with-claude/structured-outputs]
  • The tool_use approach requires an extra layer of parsing (content[0].input instead of content[0].text) and forces Claude to "call a function" rather than respond — semantically wrong for classification. output_config is the correct API.
  • The higher-level messages.parse() with a Pydantic model is convenient but requires importing Pydantic into the provider — adds coupling. Use the raw output_config dict instead, read the result from response.content[0].text, and let the existing parse_classification() in ai/utils.py handle the fallback path.

Schema to use for classification:

_CLASSIFICATION_SCHEMA = {
    "type": "object",
    "properties": {
        "assigned_topics": {"type": "array", "items": {"type": "string"}},
        "new_topic_suggestions": {"type": "array", "items": {"type": "string"}},
        "reasoning": {"type": "string"},
    },
    "required": ["assigned_topics", "new_topic_suggestions"],
    "additionalProperties": False,
}

_SUGGESTIONS_SCHEMA = {
    "type": "object",
    "properties": {
        "suggested_topics": {"type": "array", "items": {"type": "string"}},
    },
    "required": ["suggested_topics"],
    "additionalProperties": False,
}

Important limitation: output_config fails with stop_reason "refusal" or "max_tokens" — the fallback parse_classification(raw) must still be called when the response does not terminate with "end_turn".


D-07 Resolved: Client Lifecycle per Provider

Decision: Per-provider instance-level singleton — create the client once in __init__, store as self._client, reuse across all calls.

Evidence

  • AsyncOpenAI wraps an httpx.AsyncClient internally, which maintains a connection pool. Calling AsyncOpenAI(...) in _client() on every API request (current pattern) destroys the pool and forces a new TLS handshake per request. This is explicitly called out as the cause of poor concurrency in the openai-python issue tracker. [CITED: github.com/openai/openai-python/issues/1725]
  • The official recommendation is to create one AsyncOpenAI instance per process startup and reuse it. [CITED: developers.openai.com/api/reference/python]
  • AsyncAnthropic wraps httpx.AsyncClient with identical semantics. The SDK documents that both sync and async clients share the same pool-management model. [CITED: deepwiki.com/anthropics/anthropic-sdk-python/4.2-synchronous-and-asynchronous-clients]
  • httpx.AsyncClient (used for any future raw-HTTP provider) must NOT be closed between requests. It should be created once and optionally closed in an asyncio cleanup hook. [CITED: www.python-httpx.org/async/]

Breaking change warning: openai SDK 2.34.0+ raises OpenAIError: Missing credentials when api_key="". The current OllamaProvider and LMStudioProvider pass api_key="ollama" and api_key="lm-studio" respectively — these are non-empty strings and are safe. GenericOpenAIProvider must use a non-empty placeholder (e.g., "not-needed") when the user leaves the API key blank. [CITED: github.com/openai/openai-python/issues/3224]

Pattern per provider

# Pattern: store client as instance attribute, never recreate per call

class OpenAIProvider(AIProvider):
    def __init__(self, api_key: str, model: str, base_url: str | None, context_chars: int):
        self._api_key = api_key or "not-needed"   # SDK 2.34+ rejects empty string
        self._model = model
        self._base_url = base_url
        self._context_chars = context_chars
        # Singleton: created once, reused for all calls on this instance
        self._client = AsyncOpenAI(api_key=self._api_key, base_url=self._base_url)

class AnthropicProvider(AIProvider):
    def __init__(self, api_key: str, model: str, context_chars: int):
        self._api_key = api_key
        self._model = model
        self._context_chars = context_chars
        # Singleton: AsyncAnthropic manages its own httpx.AsyncClient internally
        self._client = AsyncAnthropic(api_key=self._api_key)

Why NOT FastAPI lifespan singleton: Provider instances are constructed from system_settings DB rows which can change at runtime (admin reconfigures). A module-level singleton would require re-initialization on config change. Instance-level singleton (store in self._client) is the correct level — each new ProviderConfig → new provider instance → new client. Old instance is GC'd when no longer referenced.

Async safety: AsyncOpenAI and AsyncAnthropic use asyncio-native httpx clients. They are safe to call concurrently from multiple coroutines on the same event loop (which is the Celery worker context via asyncio.run()). They are NOT thread-safe across separate OS threads. Since each Celery task call creates its own asyncio.run() event loop, each task call creates its own provider instance — no cross-task sharing occurs.


Architecture Patterns

Pattern 1: system_settings Table Design

Recommendation: One row per provider (not key-value store).

Columns:

CREATE TABLE system_settings (
    id          UUID PRIMARY KEY DEFAULT gen_random_uuid(),
    provider_id VARCHAR NOT NULL UNIQUE,   -- "openai", "anthropic", "groq", etc.
    api_key_enc TEXT,                      -- Fernet-encrypted, NULL for local providers
    base_url    TEXT,                      -- NULL means use provider SDK default
    model_name  TEXT NOT NULL DEFAULT '',  -- e.g. "gpt-4o", "claude-sonnet-4-6"
    context_chars INTEGER NOT NULL DEFAULT 8000,
    is_active   BOOLEAN NOT NULL DEFAULT FALSE,  -- only one row has is_active=TRUE at a time
    created_at  TIMESTAMPTZ NOT NULL DEFAULT now(),
    updated_at  TIMESTAMPTZ NOT NULL DEFAULT now()
);

Why one-row-per-provider over key-value:

  • Each provider has a fixed set of fields (api_key, base_url, model, context_chars) — a typed row is far better than key TEXT, value TEXT pairs that require manual assembly.
  • Enforces UNIQUE(provider_id) — no duplicate providers.
  • A single WHERE is_active = TRUE query loads the active provider config.
  • The admin UI maps naturally to a table of provider rows.

Encryption key derivation for system_settings: Unlike cloud credentials (which use user_id as the HKDF salt), system_settings API keys are global. Use provider_id as the salt:

# In services/ai_config.py — mirrors cloud_utils.py pattern exactly
def _derive_ai_settings_key(master_key: bytes, provider_id: str) -> Fernet:
    hkdf = HKDF(
        algorithm=hashes.SHA256(),
        length=32,
        salt=provider_id.encode("utf-8"),
        info=b"ai-provider-settings",  # different info= for domain separation
    )
    raw_key = hkdf.derive(master_key)
    return Fernet(base64.urlsafe_b64encode(raw_key))

Master key: Reuse settings.cloud_creds_key (already in env/config.py) — the info= parameter (b"ai-provider-settings" vs b"cloud-credentials") provides domain separation so the same master key cannot produce the same derived key for both uses.

Pattern 2: ProviderConfig Pydantic Model

# backend/ai/provider_config.py
from pydantic import BaseModel

class ProviderConfig(BaseModel):
    provider_id: str          # "openai", "anthropic", "groq", etc.
    api_key: str = ""         # decrypted at load time; empty for local providers
    base_url: str | None = None
    model: str = ""
    context_chars: int = 8000

# Named preset constructors (factories — do NOT instantiate providers here)
PROVIDER_DEFAULTS: dict[str, dict] = {
    "openai":     {"base_url": None,                    "model": "gpt-4o",              "context_chars": 120_000},
    "anthropic":  {"base_url": None,                    "model": "claude-sonnet-4-6",   "context_chars": 180_000},
    "gemini":     {"base_url": "https://generativelanguage.googleapis.com/v1beta/openai/", "model": "gemini-2.0-flash", "context_chars": 800_000},
    "groq":       {"base_url": "https://api.groq.com/openai/v1",         "model": "llama-3.3-70b-versatile", "context_chars": 128_000},
    "xai":        {"base_url": "https://api.x.ai/v1",                    "model": "grok-3-mini",  "context_chars": 128_000},
    "deepseek":   {"base_url": "https://api.deepseek.com",               "model": "deepseek-chat", "context_chars": 60_000},
    "openrouter": {"base_url": "https://openrouter.ai/api/v1",           "model": "anthropic/claude-3.5-sonnet", "context_chars": 180_000},
    "mistral":    {"base_url": "https://api.mistral.ai/v1",              "model": "mistral-large-latest", "context_chars": 128_000},
    "ollama":     {"base_url": "http://host.docker.internal:11434/v1",   "model": "llama3.2",     "context_chars": 8_000},
    "lmstudio":   {"base_url": "http://host.docker.internal:1234/v1",    "model": "gemma-4-e4b-it", "context_chars": 8_000},
}

Note: Context char values above are [ASSUMED] approximations based on known context window sizes of common models. The exact defaults are the admin's configuration responsibility — these are only startup defaults.

Pattern 3: GenericOpenAIProvider

# backend/ai/generic_openai_provider.py
from openai import AsyncOpenAI
from ai.openai_provider import OpenAIProvider

class GenericOpenAIProvider(OpenAIProvider):
    """OpenAI-compatible provider for any endpoint.

    Subclasses OpenAIProvider and enforces JSON mode on every call.
    Named presets (Groq, xAI, DeepSeek, OpenRouter, Gemini-compat, Mistral-compat,
    Ollama, LMStudio) are factory shortcuts in get_provider().
    """

    async def classify(self, document_text, existing_topics, system_prompt):
        topics_str = ", ".join(existing_topics) if existing_topics else "(none yet)"
        truncated = self._truncate(document_text)
        user_msg = f"Existing topics: [{topics_str}]\n\nDocument text:\n{truncated}"
        response = await self._client.chat.completions.create(
            model=self._model,
            max_tokens=1024,
            response_format={"type": "json_object"},   # D-01 enforcement
            messages=[
                {"role": "system", "content": system_prompt},
                {"role": "user", "content": user_msg},
            ],
        )
        raw = response.choices[0].message.content or ""
        return parse_classification(raw)   # parse_classification handles well-formed JSON too

    def _truncate(self, text: str) -> str:
        """D-13: first 60% + last 40% of context window."""
        limit = self._context_chars
        if len(text) <= limit:
            return text
        head = int(limit * 0.6)
        tail = limit - head
        return text[:head] + "\n[...truncated...]\n" + text[-tail:]

Pattern 4: Celery Retry with Exponential Backoff

The existing extract_and_classify task uses @celery_app.task(name=...). To enable self.retry(), the decorator must add bind=True.

Important: The task is a sync def that calls asyncio.run(_run(...)). The retry must happen at the sync task level, not inside _run(), because Celery's retry mechanism is synchronous (it raises Retry exception).

# backend/tasks/document_tasks.py

@celery_app.task(
    name="tasks.document_tasks.extract_and_classify",
    bind=True,
    max_retries=3,
    default_retry_delay=30,
)
def extract_and_classify(self, document_id: str) -> dict:
    """Synchronous Celery entry-point — delegates to async _run via asyncio.run."""
    try:
        return asyncio.run(_run(document_id))
    except Exception as exc:
        result = getattr(exc, "result", None)
        # Only retry on classification failures, not on extract/storage failures.
        # _run() returns a dict for extract/storage errors (non-retryable).
        # It raises only for classification exceptions.
        retry_count = self.request.retries
        countdowns = [30, 90, 270]
        countdown = countdowns[retry_count] if retry_count < len(countdowns) else 270
        raise self.retry(exc=exc, countdown=countdown)

Revised approach (cleaner): Have _run() raise ClassificationError for classification failures and return a dict for non-retryable storage/extract failures. The Celery task catches ClassificationError and calls self.retry(). On exhaustion, MaxRetriesExceededError is caught and doc.status = "classification_failed" is written.

# Cleaner separation:
@celery_app.task(
    name="tasks.document_tasks.extract_and_classify",
    bind=True,
    max_retries=3,
)
def extract_and_classify(self, document_id: str) -> dict:
    try:
        return asyncio.run(_run(document_id))
    except _ClassificationError as exc:
        retry_num = self.request.retries
        countdown = [30, 90, 270][min(retry_num, 2)]
        raise self.retry(exc=exc, countdown=countdown)
    except self.MaxRetriesExceededError:
        # All retries exhausted — write classification_failed status
        asyncio.run(_mark_classification_failed(document_id))
        return {"document_id": document_id, "status": "classification_failed"}

Document status during retries: Leave doc.status = "processing" during retries (D-10 only specifies the final state). This way the frontend shows "processing" (spinner) rather than "failed" prematurely. Only when all retries are exhausted does the status flip to "classification_failed".

Pattern 5: Re-classify Endpoint (D-11)

The existing POST /api/documents/{id}/classify in api/documents.py runs classification synchronously. It needs to be changed to re-queue the Celery task instead:

@router.post("/{doc_id}/classify")
async def reclassify_document(
    doc_id: str,
    session: AsyncSession = Depends(get_db),
    current_user: User = Depends(get_regular_user),
):
    """Re-queue the extract_and_classify Celery task for a failed document."""
    doc = await _get_owned_doc(doc_id, current_user, session)  # raises 404 on wrong owner
    doc.status = "processing"
    await session.commit()
    extract_and_classify.delay(doc_id)
    return {"document_id": doc_id, "status": "processing"}

Note: The existing endpoint already correctly does ownership check (doc.user_id != current_user.id). The change is: replace the synchronous await classifier.classify_document(...) with .delay().

Anti-Patterns to Avoid

  • Re-creating the AI client per request: The current _client(self) method in OpenAIProvider returns a new AsyncOpenAI(...) on every call. This is the pattern to remove — it prevents connection pool reuse.
  • Global MAX_AI_CHARS constant: MAX_AI_CHARS = 8_000 is defined in both openai_provider.py and anthropic_provider.py and also in classifier.py. All three must be removed and replaced by per-provider context_chars from the config.
  • Inline _settings dict in classifier.py: _settings = {"active_provider": _ai_provider, "providers": {...}} must be replaced by load_provider_config(session) returning a ProviderConfig.
  • if/elif chain in get_provider(): Replace with a registry dict for O(1) lookup.
  • Mutating shared provider state: If the admin changes config, get_provider() must return a fresh instance — never mutate self._api_key on an existing instance.

JSON Mode Compatibility Matrix

This is critical for the planner to decide which providers need the parse_classification() fallback.

Provider JSON Mode Support Notes
OpenAI (GPT-4o, etc.) Full — response_format={"type":"json_object"} [VERIFIED: developers.openai.com/api/reference/python]
Anthropic (Claude) Full — output_config.format.type="json_schema" [VERIFIED: platform.claude.com/docs/en/build-with-claude/structured-outputs]
Groq Full — response_format={"type":"json_object"} Requires schema in system prompt too [CITED: console.groq.com/docs/structured-outputs]
Mistral (compat) Full — response_format={"type":"json_object"} Must instruct model in system prompt as well [CITED: docs.mistral.ai/capabilities/structured_output/json_mode]
Gemini (via OpenAI compat) Partial — supports response_format=CalendarEvent (Pydantic/schema), NOT json_object string [CITED: ai.google.dev/gemini-api/docs/openai] — see note
xAI/Grok [ASSUMED] — likely json_object given OpenAI-compat claim Verify against xAI docs before implementing
DeepSeek [ASSUMED] — OpenAI-compat; DeepSeek-V3 supports JSON mode Verify against DeepSeek docs
OpenRouter Full — passes json_object through to underlying model [CITED: openrouter.ai/docs/guides/features/structured-outputs] — quality varies by underlying model
Ollama Partial — parameter accepted but some models ignore it [ASSUMED based on known behavior] — fallback parse_classification() is D-02
LMStudio Partial — same as Ollama (OpenAI compat, model-dependent) [ASSUMED] — fallback parse_classification() is D-02

Gemini note: The OpenAI-compat endpoint does NOT support response_format={"type": "json_object"} (string mode). It supports the structured Pydantic/schema form via response_format=CalendarEvent (using client.beta.chat.completions.parse()). For GenericOpenAIProvider which uses completions.create() with response_format={"type":"json_object"}, Gemini will likely ignore the parameter and return free-form text. [CITED: ai.google.dev/gemini-api/docs/openai]. Implication for planning: GenericOpenAIProvider should have a supports_json_mode: bool = True class-level flag, and the Gemini preset should set it to False to fall back to parse_classification().


Common Pitfalls

Pitfall 1: OpenAI SDK 2.x Empty API Key Rejection

What goes wrong: AsyncOpenAI(api_key="") raises OpenAIError: Missing credentials in SDK 2.34.0+. Affects Ollama, LMStudio, and any GenericOpenAIProvider configured without a key. Why it happens: SDK 2.34.0 added credential validation. Empty string is falsy. How to avoid: Always pass a non-empty placeholder: api_key=api_key or "not-needed". Warning signs: OpenAIError: Missing credentials during provider construction.

Pitfall 2: HKDF AlreadyFinalized if Instance is Reused

What goes wrong: Calling hkdf.derive() twice on the same HKDF object raises AlreadyFinalized. Why it happens: The cryptography library marks HKDF instances as consumed after one derivation. How to avoid: Create a fresh HKDF(...) instance on every _derive_ai_settings_key() call. Explicitly documented in cloud_utils.py — replicate the comment. Warning signs: cryptography.exceptions.AlreadyFinalized in encryption tests.

Pitfall 3: Celery bind=True + asyncio.run() Retry Scope

What goes wrong: self.retry() must be called from the sync task function, not from inside asyncio.run(_run(...)). If you raise self.retry() inside _run(), it escapes the event loop cleanup and may corrupt async state. Why it happens: asyncio.run() runs a full event loop; raising a Celery Retry exception mid-loop is not supported. How to avoid: _run() raises a custom _ClassificationError; the outer sync extract_and_classify() catches it and calls self.retry(). Warning signs: RuntimeError: Event loop is closed after retry.

Pitfall 4: is_active Dual-Write Race Condition

What goes wrong: Two concurrent admin requests both set is_active=True on different providers, resulting in two active providers. Why it happens: Read-check then write is a TOCTOU pattern. How to avoid: Use a single UPDATE system_settings SET is_active = (provider_id = $target) — updates all rows atomically. No Python-level read needed.

Pitfall 5: Anthropic output_config Schema Complexity Limit

What goes wrong: Complex classification schemas (many optional properties, union types) may hit Anthropic's grammar compilation limits: 24 optional parameters, 16 union-type parameters, 180s timeout. Why it happens: Constrained decoding compiles the schema to a grammar; complex schemas take longer. How to avoid: Keep the schema simple — assigned_topics: array[string], new_topic_suggestions: array[string], reasoning: string. All required. No unions. This is already what parse_classification() expects.

Pitfall 6: AdminAiConfigTab.vue Is Already Per-User (Wrong Scope)

What goes wrong: The existing AdminAiConfigTab.vue shows a per-user AI provider table (calling adminListUsers() and adminUpdateAiConfig(userId, ...)). Phase 7's admin AI Providers section is about global system-level provider settings, not per-user overrides. Why it happens: The existing component was built for ADMIN-05 (per-user assignment). Phase 7 adds a different concept: global system_settings. How to avoid: Add a new "Providers" sub-section in the admin AI Config tab for system-level settings. Keep the per-user assignment table. Do not confuse the two.

Pitfall 7: MAX_AI_CHARS Defined in Three Places

What goes wrong: MAX_AI_CHARS = 8_000 is defined in openai_provider.py, anthropic_provider.py, and classifier.py. Removing it from one but not the others leads to inconsistent truncation. How to avoid: A single grep-and-delete step. The planner must include this as an explicit task action.


Code Examples

Anthropic output_config Usage

# Source: platform.claude.com/docs/en/build-with-claude/structured-outputs
import anthropic

_CLASSIFICATION_SCHEMA = {
    "type": "object",
    "properties": {
        "assigned_topics": {"type": "array", "items": {"type": "string"}},
        "new_topic_suggestions": {"type": "array", "items": {"type": "string"}},
        "reasoning": {"type": "string"},
    },
    "required": ["assigned_topics", "new_topic_suggestions"],
    "additionalProperties": False,
}

class AnthropicProvider(AIProvider):
    def __init__(self, api_key: str, model: str, context_chars: int):
        self._model = model
        self._context_chars = context_chars
        self._client = anthropic.AsyncAnthropic(api_key=api_key)  # singleton

    async def classify(self, document_text, existing_topics, system_prompt):
        truncated = self._truncate(document_text)
        topics_str = ", ".join(existing_topics) if existing_topics else "(none yet)"
        user_msg = f"Existing topics: [{topics_str}]\n\nDocument text:\n{truncated}"
        response = await self._client.messages.create(
            model=self._model,
            max_tokens=1024,
            system=system_prompt,
            messages=[{"role": "user", "content": user_msg}],
            output_config={
                "format": {"type": "json_schema", "schema": _CLASSIFICATION_SCHEMA}
            },
        )
        # Fallback: if stop_reason is not "end_turn", response.content[0].text may be
        # empty or partial — parse_classification handles gracefully.
        raw = response.content[0].text if response.content else ""
        return parse_classification(raw)

Celery Exponential Backoff

# Source: docs.celeryq.dev/en/stable/userguide/tasks.html (bind + manual countdown)

class _ClassificationError(Exception):
    """Raised by _run() for retryable classification failures only."""

@celery_app.task(
    name="tasks.document_tasks.extract_and_classify",
    bind=True,
    max_retries=3,
)
def extract_and_classify(self, document_id: str) -> dict:
    """D-09: exponential backoff — 30s, 90s, 270s."""
    try:
        return asyncio.run(_run(document_id))
    except _ClassificationError as exc:
        countdowns = [30, 90, 270]
        countdown = countdowns[min(self.request.retries, 2)]
        raise self.retry(exc=exc, countdown=countdown)

Smart Truncation (D-13)

def _truncate(self, text: str) -> str:
    """First 60% + last 40% of context window (D-13)."""
    limit = self._context_chars
    if len(text) <= limit:
        return text
    head_len = int(limit * 0.6)
    tail_len = limit - head_len
    return text[:head_len] + "\n[...truncated...]\n" + text[-tail_len:]

Alembic Migration 0005

# backend/alembic/versions/0005_system_settings.py
"""Add system_settings table for AI provider configuration."""
from alembic import op
import sqlalchemy as sa

revision = "0005"
down_revision = "0004"

def upgrade():
    op.create_table(
        "system_settings",
        sa.Column("id", sa.dialects.postgresql.UUID(as_uuid=True), primary_key=True,
                  server_default=sa.text("gen_random_uuid()")),
        sa.Column("provider_id", sa.String, nullable=False),
        sa.Column("api_key_enc", sa.Text, nullable=True),
        sa.Column("base_url", sa.Text, nullable=True),
        sa.Column("model_name", sa.Text, nullable=False, server_default=""),
        sa.Column("context_chars", sa.Integer, nullable=False, server_default="8000"),
        sa.Column("is_active", sa.Boolean, nullable=False, server_default="false"),
        sa.Column("created_at", sa.TIMESTAMP(timezone=True), nullable=False,
                  server_default=sa.text("now()")),
        sa.Column("updated_at", sa.TIMESTAMP(timezone=True), nullable=False,
                  server_default=sa.text("now()")),
        sa.UniqueConstraint("provider_id", name="uq_system_settings_provider_id"),
    )

def downgrade():
    op.drop_table("system_settings")

get_provider() Registry Pattern

# backend/ai/__init__.py — O(1) lookup, no if/elif chain
from ai.provider_config import ProviderConfig
from ai.openai_provider import OpenAIProvider
from ai.anthropic_provider import AnthropicProvider
from ai.generic_openai_provider import GenericOpenAIProvider

_REGISTRY: dict[str, type] = {
    "openai":     OpenAIProvider,
    "anthropic":  AnthropicProvider,
    "gemini":     GenericOpenAIProvider,
    "groq":       GenericOpenAIProvider,
    "xai":        GenericOpenAIProvider,
    "deepseek":   GenericOpenAIProvider,
    "openrouter": GenericOpenAIProvider,
    "mistral":    GenericOpenAIProvider,
    "ollama":     GenericOpenAIProvider,
    "lmstudio":   GenericOpenAIProvider,
}

def get_provider(config: ProviderConfig) -> AIProvider:
    cls = _REGISTRY.get(config.provider_id)
    if cls is None:
        raise ValueError(f"Unknown AI provider: {config.provider_id!r}")
    return cls(
        api_key=config.api_key or "not-needed",
        model=config.model,
        base_url=config.base_url,
        context_chars=config.context_chars,
    )

Note: AnthropicProvider does not take base_url — its constructor signature will differ; handle in the factory.


Runtime State Inventory

This section is not applicable — this is a greenfield addition and refactor phase, not a rename/migration phase. No stored data keys, live service configs, or OS-registered state change names.

The one potential runtime state concern is the existing ai_provider and ai_model columns on the users table. These per-user overrides remain in place (ADMIN-05 still uses them). The new system_settings table provides the global system default that is used when a user's per-user override is NULL. The per-user fields do not change.


State of the Art

Old Approach Current Approach When Changed Impact
Anthropic tool_use for structured output output_config.format.type="json_schema" with constrained decoding ~Nov 2025 (GA by June 2026) No more JSON parsing fragility for Anthropic; model cannot produce invalid output
Celery return {"status": "classification_failed"} on exception self.retry(countdown=N, max_retries=3) then status update on exhaustion This phase Classification failures now auto-retry with backoff
_client() method creating AsyncOpenAI per call self._client = AsyncOpenAI(...) in __init__ This phase Full httpx connection pool reuse
Flat if/elif provider registry Dict-based _REGISTRY + ProviderConfig typed model This phase O(1) lookup, type safety, easy to add providers
MAX_AI_CHARS = 8_000 global constant Per-provider context_chars in ProviderConfig This phase Different providers support wildly different context sizes
Env var only AI config DB-backed system_settings with admin UI This phase Runtime reconfiguration without restarting containers

Deprecated/outdated:

  • MAX_AI_CHARS constant in openai_provider.py, anthropic_provider.py, classifier.py — remove all three.
  • The _client() method pattern in both providers — replace with __init__ singleton.
  • get_provider(settings: dict) raw dict interface — replace with get_provider(config: ProviderConfig).
  • Hardcoded if active == "anthropic" chain in __init__.py — replace with registry.

Docker Compose D-14 Status

Already done. Both backend and celery-worker services already have:

extra_hosts:
  - "host.docker.internal:host-gateway"

This was added in a prior phase. D-14 requires no docker-compose changes. The LMStudio fix is fully covered by making the base URL configurable via system_settings (D-15).


Validation Architecture

workflow.nyquist_validation = true in .planning/config.json — this section is required.

Test Framework

Property Value
Framework pytest + pytest-asyncio
Config file pytest.ini or pyproject.toml (existing)
Quick run command pytest backend/tests/test_classifier.py backend/tests/test_ai_providers.py -x
Full suite command pytest backend/tests/ -v

Phase Requirements → Test Map

Req ID Behavior Test Type Automated Command File Exists?
D-01 GenericOpenAIProvider passes response_format={"type":"json_object"} unit (mock) pytest backend/tests/test_ai_providers.py::test_generic_openai_json_mode -x No — Wave 0
D-03 AnthropicProvider passes output_config.format.type="json_schema" unit (mock) pytest backend/tests/test_ai_providers.py::test_anthropic_structured_output -x No — Wave 0
D-04 load_provider_config() reads from system_settings table integration (real DB) INTEGRATION=1 pytest backend/tests/test_ai_config.py::test_load_provider_config -x No — Wave 0
D-05 system_settings API keys encrypt/decrypt correctly unit pytest backend/tests/test_ai_config.py::test_api_key_encrypt_decrypt -x No — Wave 0
D-05 Admin GET /api/admin/ai-config never returns api_key_enc integration pytest backend/tests/test_admin_ai_config.py::test_get_never_returns_key -x No — Wave 0
D-06 get_provider() accepts ProviderConfig, not raw dict unit pytest backend/tests/test_ai_providers.py::test_get_provider_typed -x No — Wave 0
D-07 OpenAIProvider._client is not recreated on second call unit pytest backend/tests/test_ai_providers.py::test_client_singleton -x No — Wave 0
D-09 extract_and_classify retries with 30s/90s/270s backoff unit (mock) pytest backend/tests/test_document_tasks.py::test_retry_backoff -x No — Wave 0
D-10 After 3 retries doc.status = "classification_failed" unit (mock) pytest backend/tests/test_document_tasks.py::test_exhaustion_sets_failed_status -x No — Wave 0
D-12 context_chars respected per provider unit pytest backend/tests/test_ai_providers.py::test_context_chars_truncation -x No — Wave 0
D-13 Truncation is 60% head + 40% tail unit pytest backend/tests/test_ai_providers.py::test_smart_truncation -x No — Wave 0
D-11 POST /api/documents/{id}/classify re-queues Celery integration pytest backend/tests/test_documents.py::test_reclassify_requeues_celery -x No — Wave 0
existing parse_classification/parse_suggestions robustness unit pytest backend/tests/test_classifier.py -x Yes

Sampling Rate

  • Per task commit: pytest backend/tests/test_classifier.py backend/tests/test_ai_providers.py -x
  • Per wave merge: pytest backend/tests/ -v
  • Phase gate: Full suite green before /gsd:verify-work

Wave 0 Gaps

  • backend/tests/test_ai_providers.py — covers D-01, D-03, D-06, D-07, D-12, D-13
  • backend/tests/test_ai_config.py — covers D-04, D-05 (encryption round-trip + DB read)
  • backend/tests/test_admin_ai_config.py — covers admin endpoint security (never returns key)
  • backend/tests/test_document_tasks.py (new tests in existing file) — covers D-09, D-10

Note: test_classifier.py already exists and covers parse_classification() robustness — keep as-is.


Security Domain

Applicable ASVS Categories

ASVS Category Applies Standard Control
V2 Authentication No
V3 Session Management No
V4 Access Control Yes get_current_admin on all /api/admin/ai-config endpoints
V5 Input Validation Yes Pydantic ProviderConfig validates all admin inputs
V6 Cryptography Yes HKDF/AES-GCM (same pattern as cloud credentials)

Known Threat Patterns

Pattern STRIDE Standard Mitigation
Admin endpoint returns api_key_enc Information Disclosure _ai_config_to_dict() whitelist — never include api_key_enc field
Prompt injection via base_url or model_name Tampering Pydantic validation; base_url validated against URL format; model_name is opaque string passed to provider
HKDF key reuse across domains Elevation of Privilege info=b"ai-provider-settings"info=b"cloud-credentials" — domain separation
Non-admin user reads system AI config Unauthorized Access Admin-only endpoints (get_current_admin dep)
Classification of attacker-controlled documents Tampering Document content is user's own; provider calls are server-side; no SSRF from provider base_url (it's admin-configured, not user-supplied)

Admin endpoint invariant: GET /api/admin/ai-config and PUT /api/admin/ai-config must never return api_key_enc. Follow the _user_to_dict() whitelist pattern from admin.py — create _ai_config_to_dict(row) that includes provider_id, base_url, model_name, context_chars, is_active, updated_at but excludes api_key_enc. The API key field in the admin UI is write-only (masked input; never pre-filled from the server).


Environment Availability

Dependency Required By Available Version Fallback
PostgreSQL system_settings table Yes (Docker) 17-alpine
Redis Celery retry queue Yes (Docker) 7-alpine
Python openai SDK ≥1.30 GenericOpenAIProvider Yes (requirements.txt) Latest: 2.40.0
Python anthropic SDK ≥0.26 AnthropicProvider Yes (requirements.txt) Latest: 0.105.2
Python celery ≥5.5.0 bind=True retry Yes (requirements.txt) Latest: 5.6.3
docker-compose extra_hosts LMStudio/Ollama host access Already present in docker-compose.yml

Missing dependencies with no fallback: None.


Assumptions Log

# Claim Section Risk if Wrong
A1 xAI/Grok API supports response_format={"type":"json_object"} JSON Compatibility Matrix Planner must add supports_json_mode=False preset for xAI if not supported; fallback to parse_classification()
A2 DeepSeek supports response_format={"type":"json_object"} JSON Compatibility Matrix Same as A1 — fallback is already designed in D-02
A3 Context char defaults in PROVIDER_DEFAULTS table Pattern 2 Admin will reconfigure; only startup defaults; no functional risk
A4 Ollama/LMStudio: response_format param is accepted but may be ignored by some models JSON Compatibility Matrix Already handled by D-02 fallback; no implementation risk
A5 anthropic SDK version ≥0.95 is installed (output_config support added around v0.95) D-03 section If older version: output_config kwarg raises TypeError; bump requirement to >=0.95.0

Note on A5: The requirements.txt pins anthropic>=0.26. The output_config parameter was introduced significantly after that. The planner must update the minimum pin to anthropic>=0.95.0 (or preferably the current latest 0.105.2).


Open Questions (RESOLVED)

  1. xAI/Grok and DeepSeek JSON mode verification

    • What we know: Both claim OpenAI-compat; industry pattern suggests they support json_object
    • What's unclear: Neither was directly verified against official docs in this research session
    • Recommendation: Planner should include a Wave 0 task to spot-check against official xAI and DeepSeek API docs, or treat both as [ASSUMED] and add supports_json_mode flag that defaults to True but can be overridden
    • RESOLVED: treated as [ASSUMED]; supports_json_mode flag is the mitigation — planner accepts the risk. xAI and DeepSeek are configured with supports_json_mode=True in Plan 02; if a future user reports JSON-shape failures, flipping the SUPPORTS_JSON_MODE entry to False routes that provider through parse_classification() (D-02 fallback path). No blocking action required for Phase 7 execution.
  2. anthropic SDK minimum version for output_config

    • What we know: output_config is confirmed in current SDK (0.105.2) with no beta headers needed
    • What's unclear: The exact version when output_config was introduced (the old output_format/beta-header path still works)
    • Recommendation: Bump requirements.txt to anthropic>=0.95.0 as part of Wave 1; if output_format param (old alias) is still working, the older SDK might work too — but it's safer to pin to a known good version
    • RESOLVED: treated as [ASSUMED]; supports_json_mode flag is the mitigation — planner accepts the risk. Plan 02 pins anthropic>=0.95.0 in requirements.txt; if a stricter floor is needed it can be raised in a follow-up phase. The fallback path (parse_classification) handles any output_config regression.
  3. Gemini-compat: use json_object or fall back?

    • What we know: The OpenAI-compat endpoint does NOT document json_object string mode; it supports Pydantic schema via client.beta.chat.completions.parse()
    • What's unclear: Whether json_object is silently ignored or raises an error
    • Recommendation: Implement Gemini preset with supports_json_mode=False; GenericOpenAIProvider.classify() skips response_format for these presets and relies on parse_classification() fallback
    • RESOLVED: treated as [ASSUMED]; supports_json_mode flag is the mitigation — planner accepts the risk. Plan 02 SUPPORTS_JSON_MODE["gemini"] = False routes Gemini through the parse_classification() fallback path. Test test_gemini_fallback_to_parse_classification in Plan 02 Task 4 enforces this contract.

Sources

Primary (HIGH confidence)

Secondary (MEDIUM confidence)

Tertiary (LOW confidence)

  • Ollama/LMStudio JSON mode behaviour — [ASSUMED] from known community patterns; not verified against official Ollama API docs in this session
  • xAI/Grok json_object support — [ASSUMED]
  • DeepSeek json_object support — [ASSUMED]

Metadata

Confidence breakdown:

  • Standard stack: HIGH — all packages already in requirements.txt, versions verified on PyPI
  • Anthropic structured output (D-03): HIGH — verified against official Anthropic docs
  • Client lifecycle (D-07): HIGH — verified against openai-python issues + Anthropic SDK docs
  • JSON mode matrix: MEDIUM — Groq/Mistral/OpenRouter verified; xAI/DeepSeek/Gemini partially assumed
  • Celery retry pattern: HIGH — verified against official Celery 5 docs
  • system_settings design: HIGH — based on established cloud_utils.py pattern in codebase
  • Pitfalls: HIGH — openai SDK empty key bug directly cited from issue tracker

Research date: 2026-06-02 Valid until: 2026-09-01 (stable APIs; Anthropic structured output is GA, no longer beta)