5 waves: system_settings DB + HKDF encryption (01), ProviderConfig + GenericOpenAIProvider + singleton client fix (02), Anthropic output_config + classifier wiring (03), Celery retry 30/90/270s + re-queue endpoint (04), admin AI panel + DocumentCard badge + human checkpoint (05). All 18 decisions D-01..D-18 covered. Plan checker passed after 1 revision round (4 blockers fixed: D-02/D-14 coverage, D-11 Vitest tests, Plan 05 files_modified). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
55 KiB
Phase 7: Redo and Optimize LLM Integration - Research
Researched: 2026-06-02 Domain: AI provider abstraction layer, Celery retry, structured output, DB-driven config Confidence: HIGH
<user_constraints>
User Constraints (from CONTEXT.md)
Locked Decisions
- D-01: All OpenAI-compatible providers (OpenAI, Ollama, LMStudio, Groq, Gemini-compat, Mistral-compat, xAI/Grok, DeepSeek, OpenRouter) must pass
response_format={"type": "json_object"}in everyclassify()andsuggest_topics()call. - D-02: Keep
parse_classification()andparse_suggestions()inai/utils.pyas a last-resort fallback only — they are called when a provider does not supportresponse_format(e.g., some local Ollama models that ignore the parameter). - D-04: All AI provider settings (API keys, base URLs, model names, active provider) are stored in a new
system_settingsDB table. Env vars serve as startup defaults that populate the table on first boot, but the DB record is authoritative at runtime. - D-05: API keys stored in
system_settingsmust be encrypted at rest — use the same HKDF/AES-GCM pattern already established for cloud credentials inbackend/storage/cloud_utils.py. - D-06:
get_provider()accepts typed PydanticProviderConfigmodels (not raw dicts). The classifier assembles aProviderConfigfrom the DB record. - D-08: Admin panel gains an "AI Providers" configuration section where admins can set the active provider, configure API keys/base URLs/models, and test connectivity via the existing
health_check()method. - D-09: The
extract_and_classifyCelery task uses Celery's built-in retry mechanism (self.retry(exc=…, countdown=…, max_retries=3)) with exponential backoff: first retry at 30 s, second at 90 s, third at 270 s. - D-10: After all retries are exhausted the document status remains
classification_failed. No silent failure. - D-11: The document card shows a "Classification failed" badge when
status == "classification_failed". A "Re-analyze" button callsPOST /api/documents/{id}/classifyto re-queue the Celery task. - D-12:
MAX_AI_CHARSis replaced by a per-providercontext_charsvalue declared in each provider's config (orProviderConfigdefault). The global constant is removed. - D-13: Smart truncation strategy: take the first 60% of the per-provider limit + the last 40%.
- D-14: Add
extra_hosts: ["host.docker.internal:host-gateway"]to thebackendandcelery-workerservices indocker-compose.yml. - D-15: LMStudio base URL and model are configurable via the admin AI Providers panel.
- D-16: Introduce a
GenericOpenAIProvider(base_url, api_key, model, context_chars)class inbackend/ai/generic_openai_provider.py. It subclassesOpenAIProviderand adds JSON-mode enforcement. - D-17: Providers to add: Google Gemini (via OpenAI-compat endpoint), Mistral (via OpenAI-compat), Groq, xAI/Grok, DeepSeek, OpenRouter — all via
GenericOpenAIProvider. - D-18: Mistral uses the OpenAI-compat endpoint (not the native
mistralaiSDK).
Claude's Discretion
- D-03 (Anthropic output): Researcher determines the most reliable pattern for structured JSON from Claude models consistent with D-01.
- D-07 (Client lifecycle): Researcher documents optimal client instantiation per provider (singleton, per-request, pooled) clearly enough that adding a new provider requires no re-research.
Deferred Ideas (OUT OF SCOPE)
- Chunk + merge multi-call classification (too expensive per-API call in this phase)
- Native
mistralaiSDK (OpenAI-compat is sufficient) - Token counting with
tiktokenor provider tokenizers - Streaming classification
</user_constraints>
Summary
Phase 7 is a focused refactor of an already-functional AI layer. The goal is reliability and extensibility: move provider settings from env vars into a DB table (system_settings), introduce a GenericOpenAIProvider that covers all OpenAI-compatible APIs in one class, enforce JSON-mode structured output across all providers, add Celery retry with exponential backoff for classification failures, and surface a "Re-analyze" button in the UI when classification fails.
The research resolves the two open design questions: (1) Anthropic now supports native output_config.format.type="json_schema" on all current Claude models with constrained decoding — no tool_use workaround is needed; (2) the optimal client lifecycle pattern is a per-provider module-level singleton that is constructed once at provider instantiation and stored as self._client, not recreated per API call.
A critical codebase finding: extra_hosts: ["host.docker.internal:host-gateway"] is already present on both backend and celery-worker in docker-compose.yml — D-14 is already done. Similarly, POST /api/documents/{id}/classify already exists in api/documents.py and classifyDocument() is already in the frontend API client. The phase needs to change what happens inside that endpoint (re-queue Celery instead of running synchronously) and update the frontend badge logic, not add new routing from scratch.
Primary recommendation: Implement in five waves: (1) Alembic migration + system_settings model + encryption helpers, (2) ProviderConfig Pydantic models + GenericOpenAIProvider + refactored provider constructors, (3) Anthropic output_config structured output + classifier refactor with smart truncation, (4) Celery retry harness + re-classify endpoint to re-queue, (5) Admin AI Config panel UI + DocumentCard badge + tests.
Architectural Responsibility Map
| Capability | Primary Tier | Secondary Tier | Rationale |
|---|---|---|---|
| system_settings read/write | Database (PostgreSQL ORM) | API / Backend | Settings are global admin config, live in DB |
| API key encryption/decryption | API / Backend (services layer) | — | Crypto must never happen in the DB or frontend |
| Provider factory / get_provider() | API / Backend (ai/__init__.py) |
— | Instantiation is a backend concern |
| JSON structured output enforcement | API / Backend (each provider class) | — | Each provider's classify() sets response_format |
| Context window truncation | API / Backend (each provider class) | — | Done before calling the provider API |
| Celery retry logic | Celery worker (tasks/document_tasks.py) |
— | Task retry is a worker-level concern |
| Re-classify trigger | API / Backend endpoint | Frontend (button) | Backend re-queues; frontend emits the call |
| Admin AI config UI | Frontend (AdminAiConfigTab.vue) | Admin API endpoint | UI displays/saves; backend enforces |
| Classification failed badge | Frontend (DocumentCard.vue) | — | UI reads doc.status == "classification_failed" |
Standard Stack
Core (no new packages — all already in requirements.txt)
| Library | Version (pinned min) | Purpose | Why Standard |
|---|---|---|---|
openai |
>=1.30 (latest: 2.40.0) |
AsyncOpenAI for OpenAI + all compat providers | Official SDK; manages httpx connection pool |
anthropic |
>=0.26 (latest: 0.105.2) |
AsyncAnthropic for native Claude output | Official SDK; supports output_config.format |
celery[redis] |
>=5.5.0 (latest: 5.6.3) |
Task retry with exponential backoff | Already used for document processing |
cryptography |
>=41.0.0 |
HKDF/AES-GCM for system_settings API key encryption |
Same pattern as cloud credentials |
pydantic |
>=2.0 |
ProviderConfig typed models |
Already used everywhere |
sqlalchemy[asyncio] |
>=2.0.49 |
SystemSettings ORM model |
Already used everywhere |
httpx |
>=0.27 (latest: 0.28.1) |
Transport layer for all async HTTP clients | Used internally by both openai and anthropic SDKs |
No new packages need to be added to requirements.txt for this phase. All required libraries are already present.
Alternatives Considered
| Instead of | Could Use | Tradeoff |
|---|---|---|
openai SDK for compat providers |
httpx directly |
openai SDK manages connection pooling, retries, and auth headers automatically — less code |
output_config.format for Anthropic |
tool_use JSON schema |
output_config is simpler, no tool_call parsing needed, and uses constrained decoding |
Per-provider context_chars int |
tiktoken token counts |
Char approximation is 4x cheaper in CPU and requires no extra dependency; deferred in CONTEXT.md |
Package Legitimacy Audit
No new packages are introduced in this phase. All libraries (openai, anthropic, celery, cryptography, pydantic, sqlalchemy, httpx) are long-established, present in requirements.txt, and well-known to the ecosystem.
| Package | Registry | Age | Status |
|---|---|---|---|
openai |
PyPI | ~5 yrs | Established — official OpenAI Python SDK |
anthropic |
PyPI | ~3 yrs | Established — official Anthropic Python SDK |
celery |
PyPI | ~15 yrs | Established |
cryptography |
PyPI | ~10 yrs | Established |
Packages removed due to slopcheck verdict: none — no new packages.
Architecture Patterns
System Architecture Diagram
[Admin Browser]
│ PUT /api/admin/ai-config
▼
[Admin API (admin.py)]
│ HKDF-encrypt(api_key) write to system_settings table
▼
[PostgreSQL: system_settings table]
│ load_provider_config(session) reads on each classify call
▼
[classifier.py — classify_document()]
│ builds ProviderConfig calls get_provider(config)
▼
[ai/__init__.py — get_provider(config: ProviderConfig)]
│ registry lookup construct AnthropicProvider | GenericOpenAIProvider
▼
[Provider instance (singleton _client)]
│ classify() / suggest_topics() send request to LLM API
▼
[External LLM API]
│ response JSON (or fallback parse_classification())
▼
[classifier.py] — persist topics to DB
▲
[Celery task: extract_and_classify (bind=True)]
│ on exception: self.retry(countdown=30|90|270, max_retries=3)
│ on exhaustion: doc.status = "classification_failed"
▼
[Document: status field]
│ polled by frontend
▼
[DocumentCard.vue]
└── classification_failed badge + Re-analyze button
│ POST /api/documents/{id}/classify
▼
[documents.py — re-queues Celery task]
Recommended Project Structure (additions only)
backend/
├── ai/
│ ├── base.py # AIProvider ABC (no change)
│ ├── utils.py # parse_classification/parse_suggestions (no change)
│ ├── __init__.py # get_provider(config: ProviderConfig) — refactored
│ ├── provider_config.py # NEW: ProviderConfig Pydantic model + presets
│ ├── generic_openai_provider.py # NEW: GenericOpenAIProvider(OpenAIProvider)
│ ├── openai_provider.py # ADD context_chars, singleton _client, json_mode
│ ├── anthropic_provider.py # ADD output_config structured output, singleton _client
│ ├── ollama_provider.py # UPDATE: pass context_chars
│ └── lmstudio_provider.py # UPDATE: pass context_chars
├── db/
│ └── models.py # ADD SystemSettings ORM model
├── alembic/versions/
│ └── 0005_system_settings.py # NEW: migration for system_settings table
├── services/
│ └── ai_config.py # NEW: load_provider_config(session) helper
└── tasks/
└── document_tasks.py # REFACTOR: bind=True, self.retry() exponential backoff
frontend/src/
├── components/
│ ├── admin/
│ │ └── AdminAiConfigTab.vue # REPLACE: global provider config (not per-user)
│ └── documents/
│ └── DocumentCard.vue # ADD: classification_failed badge + Re-analyze button
└── api/
└── client.js # ADD: reclassifyDocument(), getAiConfig(), saveAiConfig()
Don't Hand-Roll
| Problem | Don't Build | Use Instead | Why |
|---|---|---|---|
| JSON schema enforcement on LLM output | Custom regex/JSON repair | response_format={"type":"json_object"} (OpenAI compat) or output_config.format (Anthropic) |
SDK-level constraint; cannot produce invalid JSON with constrained decoding |
| Exponential backoff with jitter | Manual sleep loop | self.retry(countdown=N, max_retries=3) in Celery |
Celery handles scheduling, worker redelivery, and failure state atomically |
| Per-provider connection pools | Custom httpx session management | Module-level AsyncOpenAI / AsyncAnthropic singletons |
Both SDKs manage httpx connection pools internally; re-instantiation destroys pool reuse |
| HKDF key derivation | AES with hardcoded key | _derive_fernet_key(master_key, provider_id) from storage/cloud_utils.py |
Pattern already battle-tested in Phase 5; reuse verbatim |
| Provider capability detection | Runtime feature probe | Static supports_json_mode: bool flag per preset |
Avoids extra API calls; Ollama/LMStudio's fallback is already handled in D-02 |
Key insight: The entire OpenAI-compat ecosystem (Groq, xAI, DeepSeek, OpenRouter, Gemini-compat, Mistral-compat) is handled by a single GenericOpenAIProvider with different base_url defaults. The JSON-mode parameter travels through unchanged. Do not implement separate classes for each.
D-03 Resolved: Anthropic Structured Output Strategy
Decision: Use output_config={"format": {"type": "json_schema", "schema": {...}}} in messages.create().
Rationale:
- The Anthropic Python SDK (v0.95+) supports
output_config.format.type = "json_schema"with constrained decoding — the model is physically prevented from producing invalid JSON matching the schema. [VERIFIED: platform.claude.com/docs/en/build-with-claude/structured-outputs] - Supported on all current Claude models: Claude Haiku 4.5, Sonnet 4.5, Sonnet 4.6, Opus 4.5, Opus 4.6, Opus 4.7, Opus 4.8. [VERIFIED: platform.claude.com/docs/en/build-with-claude/structured-outputs]
- No beta headers are required as of SDK ≥0.95 (the old
structured-outputs-2025-11-13header is still accepted for backward compat but no longer needed). [VERIFIED: platform.claude.com/docs/en/build-with-claude/structured-outputs] - The
tool_useapproach requires an extra layer of parsing (content[0].inputinstead ofcontent[0].text) and forces Claude to "call a function" rather than respond — semantically wrong for classification.output_configis the correct API. - The higher-level
messages.parse()with a Pydantic model is convenient but requires importing Pydantic into the provider — adds coupling. Use the rawoutput_configdict instead, read the result fromresponse.content[0].text, and let the existingparse_classification()inai/utils.pyhandle the fallback path.
Schema to use for classification:
_CLASSIFICATION_SCHEMA = {
"type": "object",
"properties": {
"assigned_topics": {"type": "array", "items": {"type": "string"}},
"new_topic_suggestions": {"type": "array", "items": {"type": "string"}},
"reasoning": {"type": "string"},
},
"required": ["assigned_topics", "new_topic_suggestions"],
"additionalProperties": False,
}
_SUGGESTIONS_SCHEMA = {
"type": "object",
"properties": {
"suggested_topics": {"type": "array", "items": {"type": "string"}},
},
"required": ["suggested_topics"],
"additionalProperties": False,
}
Important limitation: output_config fails with stop_reason "refusal" or "max_tokens" — the fallback parse_classification(raw) must still be called when the response does not terminate with "end_turn".
D-07 Resolved: Client Lifecycle per Provider
Decision: Per-provider instance-level singleton — create the client once in __init__, store as self._client, reuse across all calls.
Evidence
AsyncOpenAIwraps anhttpx.AsyncClientinternally, which maintains a connection pool. CallingAsyncOpenAI(...)in_client()on every API request (current pattern) destroys the pool and forces a new TLS handshake per request. This is explicitly called out as the cause of poor concurrency in the openai-python issue tracker. [CITED: github.com/openai/openai-python/issues/1725]- The official recommendation is to create one
AsyncOpenAIinstance per process startup and reuse it. [CITED: developers.openai.com/api/reference/python] AsyncAnthropicwrapshttpx.AsyncClientwith identical semantics. The SDK documents that both sync and async clients share the same pool-management model. [CITED: deepwiki.com/anthropics/anthropic-sdk-python/4.2-synchronous-and-asynchronous-clients]httpx.AsyncClient(used for any future raw-HTTP provider) must NOT be closed between requests. It should be created once and optionally closed in anasynciocleanup hook. [CITED: www.python-httpx.org/async/]
Breaking change warning: openai SDK 2.34.0+ raises OpenAIError: Missing credentials when api_key="". The current OllamaProvider and LMStudioProvider pass api_key="ollama" and api_key="lm-studio" respectively — these are non-empty strings and are safe. GenericOpenAIProvider must use a non-empty placeholder (e.g., "not-needed") when the user leaves the API key blank. [CITED: github.com/openai/openai-python/issues/3224]
Pattern per provider
# Pattern: store client as instance attribute, never recreate per call
class OpenAIProvider(AIProvider):
def __init__(self, api_key: str, model: str, base_url: str | None, context_chars: int):
self._api_key = api_key or "not-needed" # SDK 2.34+ rejects empty string
self._model = model
self._base_url = base_url
self._context_chars = context_chars
# Singleton: created once, reused for all calls on this instance
self._client = AsyncOpenAI(api_key=self._api_key, base_url=self._base_url)
class AnthropicProvider(AIProvider):
def __init__(self, api_key: str, model: str, context_chars: int):
self._api_key = api_key
self._model = model
self._context_chars = context_chars
# Singleton: AsyncAnthropic manages its own httpx.AsyncClient internally
self._client = AsyncAnthropic(api_key=self._api_key)
Why NOT FastAPI lifespan singleton: Provider instances are constructed from system_settings DB rows which can change at runtime (admin reconfigures). A module-level singleton would require re-initialization on config change. Instance-level singleton (store in self._client) is the correct level — each new ProviderConfig → new provider instance → new client. Old instance is GC'd when no longer referenced.
Async safety: AsyncOpenAI and AsyncAnthropic use asyncio-native httpx clients. They are safe to call concurrently from multiple coroutines on the same event loop (which is the Celery worker context via asyncio.run()). They are NOT thread-safe across separate OS threads. Since each Celery task call creates its own asyncio.run() event loop, each task call creates its own provider instance — no cross-task sharing occurs.
Architecture Patterns
Pattern 1: system_settings Table Design
Recommendation: One row per provider (not key-value store).
Columns:
CREATE TABLE system_settings (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
provider_id VARCHAR NOT NULL UNIQUE, -- "openai", "anthropic", "groq", etc.
api_key_enc TEXT, -- Fernet-encrypted, NULL for local providers
base_url TEXT, -- NULL means use provider SDK default
model_name TEXT NOT NULL DEFAULT '', -- e.g. "gpt-4o", "claude-sonnet-4-6"
context_chars INTEGER NOT NULL DEFAULT 8000,
is_active BOOLEAN NOT NULL DEFAULT FALSE, -- only one row has is_active=TRUE at a time
created_at TIMESTAMPTZ NOT NULL DEFAULT now(),
updated_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
Why one-row-per-provider over key-value:
- Each provider has a fixed set of fields (api_key, base_url, model, context_chars) — a typed row is far better than
key TEXT, value TEXTpairs that require manual assembly. - Enforces
UNIQUE(provider_id)— no duplicate providers. - A single
WHERE is_active = TRUEquery loads the active provider config. - The admin UI maps naturally to a table of provider rows.
Encryption key derivation for system_settings: Unlike cloud credentials (which use user_id as the HKDF salt), system_settings API keys are global. Use provider_id as the salt:
# In services/ai_config.py — mirrors cloud_utils.py pattern exactly
def _derive_ai_settings_key(master_key: bytes, provider_id: str) -> Fernet:
hkdf = HKDF(
algorithm=hashes.SHA256(),
length=32,
salt=provider_id.encode("utf-8"),
info=b"ai-provider-settings", # different info= for domain separation
)
raw_key = hkdf.derive(master_key)
return Fernet(base64.urlsafe_b64encode(raw_key))
Master key: Reuse settings.cloud_creds_key (already in env/config.py) — the info= parameter (b"ai-provider-settings" vs b"cloud-credentials") provides domain separation so the same master key cannot produce the same derived key for both uses.
Pattern 2: ProviderConfig Pydantic Model
# backend/ai/provider_config.py
from pydantic import BaseModel
class ProviderConfig(BaseModel):
provider_id: str # "openai", "anthropic", "groq", etc.
api_key: str = "" # decrypted at load time; empty for local providers
base_url: str | None = None
model: str = ""
context_chars: int = 8000
# Named preset constructors (factories — do NOT instantiate providers here)
PROVIDER_DEFAULTS: dict[str, dict] = {
"openai": {"base_url": None, "model": "gpt-4o", "context_chars": 120_000},
"anthropic": {"base_url": None, "model": "claude-sonnet-4-6", "context_chars": 180_000},
"gemini": {"base_url": "https://generativelanguage.googleapis.com/v1beta/openai/", "model": "gemini-2.0-flash", "context_chars": 800_000},
"groq": {"base_url": "https://api.groq.com/openai/v1", "model": "llama-3.3-70b-versatile", "context_chars": 128_000},
"xai": {"base_url": "https://api.x.ai/v1", "model": "grok-3-mini", "context_chars": 128_000},
"deepseek": {"base_url": "https://api.deepseek.com", "model": "deepseek-chat", "context_chars": 60_000},
"openrouter": {"base_url": "https://openrouter.ai/api/v1", "model": "anthropic/claude-3.5-sonnet", "context_chars": 180_000},
"mistral": {"base_url": "https://api.mistral.ai/v1", "model": "mistral-large-latest", "context_chars": 128_000},
"ollama": {"base_url": "http://host.docker.internal:11434/v1", "model": "llama3.2", "context_chars": 8_000},
"lmstudio": {"base_url": "http://host.docker.internal:1234/v1", "model": "gemma-4-e4b-it", "context_chars": 8_000},
}
Note: Context char values above are [ASSUMED] approximations based on known context window sizes of common models. The exact defaults are the admin's configuration responsibility — these are only startup defaults.
Pattern 3: GenericOpenAIProvider
# backend/ai/generic_openai_provider.py
from openai import AsyncOpenAI
from ai.openai_provider import OpenAIProvider
class GenericOpenAIProvider(OpenAIProvider):
"""OpenAI-compatible provider for any endpoint.
Subclasses OpenAIProvider and enforces JSON mode on every call.
Named presets (Groq, xAI, DeepSeek, OpenRouter, Gemini-compat, Mistral-compat,
Ollama, LMStudio) are factory shortcuts in get_provider().
"""
async def classify(self, document_text, existing_topics, system_prompt):
topics_str = ", ".join(existing_topics) if existing_topics else "(none yet)"
truncated = self._truncate(document_text)
user_msg = f"Existing topics: [{topics_str}]\n\nDocument text:\n{truncated}"
response = await self._client.chat.completions.create(
model=self._model,
max_tokens=1024,
response_format={"type": "json_object"}, # D-01 enforcement
messages=[
{"role": "system", "content": system_prompt},
{"role": "user", "content": user_msg},
],
)
raw = response.choices[0].message.content or ""
return parse_classification(raw) # parse_classification handles well-formed JSON too
def _truncate(self, text: str) -> str:
"""D-13: first 60% + last 40% of context window."""
limit = self._context_chars
if len(text) <= limit:
return text
head = int(limit * 0.6)
tail = limit - head
return text[:head] + "\n[...truncated...]\n" + text[-tail:]
Pattern 4: Celery Retry with Exponential Backoff
The existing extract_and_classify task uses @celery_app.task(name=...). To enable self.retry(), the decorator must add bind=True.
Important: The task is a sync def that calls asyncio.run(_run(...)). The retry must happen at the sync task level, not inside _run(), because Celery's retry mechanism is synchronous (it raises Retry exception).
# backend/tasks/document_tasks.py
@celery_app.task(
name="tasks.document_tasks.extract_and_classify",
bind=True,
max_retries=3,
default_retry_delay=30,
)
def extract_and_classify(self, document_id: str) -> dict:
"""Synchronous Celery entry-point — delegates to async _run via asyncio.run."""
try:
return asyncio.run(_run(document_id))
except Exception as exc:
result = getattr(exc, "result", None)
# Only retry on classification failures, not on extract/storage failures.
# _run() returns a dict for extract/storage errors (non-retryable).
# It raises only for classification exceptions.
retry_count = self.request.retries
countdowns = [30, 90, 270]
countdown = countdowns[retry_count] if retry_count < len(countdowns) else 270
raise self.retry(exc=exc, countdown=countdown)
Revised approach (cleaner): Have _run() raise ClassificationError for classification failures and return a dict for non-retryable storage/extract failures. The Celery task catches ClassificationError and calls self.retry(). On exhaustion, MaxRetriesExceededError is caught and doc.status = "classification_failed" is written.
# Cleaner separation:
@celery_app.task(
name="tasks.document_tasks.extract_and_classify",
bind=True,
max_retries=3,
)
def extract_and_classify(self, document_id: str) -> dict:
try:
return asyncio.run(_run(document_id))
except _ClassificationError as exc:
retry_num = self.request.retries
countdown = [30, 90, 270][min(retry_num, 2)]
raise self.retry(exc=exc, countdown=countdown)
except self.MaxRetriesExceededError:
# All retries exhausted — write classification_failed status
asyncio.run(_mark_classification_failed(document_id))
return {"document_id": document_id, "status": "classification_failed"}
Document status during retries: Leave doc.status = "processing" during retries (D-10 only specifies the final state). This way the frontend shows "processing" (spinner) rather than "failed" prematurely. Only when all retries are exhausted does the status flip to "classification_failed".
Pattern 5: Re-classify Endpoint (D-11)
The existing POST /api/documents/{id}/classify in api/documents.py runs classification synchronously. It needs to be changed to re-queue the Celery task instead:
@router.post("/{doc_id}/classify")
async def reclassify_document(
doc_id: str,
session: AsyncSession = Depends(get_db),
current_user: User = Depends(get_regular_user),
):
"""Re-queue the extract_and_classify Celery task for a failed document."""
doc = await _get_owned_doc(doc_id, current_user, session) # raises 404 on wrong owner
doc.status = "processing"
await session.commit()
extract_and_classify.delay(doc_id)
return {"document_id": doc_id, "status": "processing"}
Note: The existing endpoint already correctly does ownership check (doc.user_id != current_user.id). The change is: replace the synchronous await classifier.classify_document(...) with .delay().
Anti-Patterns to Avoid
- Re-creating the AI client per request: The current
_client(self)method inOpenAIProviderreturns a newAsyncOpenAI(...)on every call. This is the pattern to remove — it prevents connection pool reuse. - Global MAX_AI_CHARS constant:
MAX_AI_CHARS = 8_000is defined in bothopenai_provider.pyandanthropic_provider.pyand also inclassifier.py. All three must be removed and replaced by per-providercontext_charsfrom the config. - Inline
_settingsdict in classifier.py:_settings = {"active_provider": _ai_provider, "providers": {...}}must be replaced byload_provider_config(session)returning aProviderConfig. if/elifchain inget_provider(): Replace with a registry dict for O(1) lookup.- Mutating shared provider state: If the admin changes config,
get_provider()must return a fresh instance — never mutateself._api_keyon an existing instance.
JSON Mode Compatibility Matrix
This is critical for the planner to decide which providers need the parse_classification() fallback.
| Provider | JSON Mode Support | Notes |
|---|---|---|
| OpenAI (GPT-4o, etc.) | Full — response_format={"type":"json_object"} |
[VERIFIED: developers.openai.com/api/reference/python] |
| Anthropic (Claude) | Full — output_config.format.type="json_schema" |
[VERIFIED: platform.claude.com/docs/en/build-with-claude/structured-outputs] |
| Groq | Full — response_format={"type":"json_object"} |
Requires schema in system prompt too [CITED: console.groq.com/docs/structured-outputs] |
| Mistral (compat) | Full — response_format={"type":"json_object"} |
Must instruct model in system prompt as well [CITED: docs.mistral.ai/capabilities/structured_output/json_mode] |
| Gemini (via OpenAI compat) | Partial — supports response_format=CalendarEvent (Pydantic/schema), NOT json_object string |
[CITED: ai.google.dev/gemini-api/docs/openai] — see note |
| xAI/Grok | [ASSUMED] — likely json_object given OpenAI-compat claim |
Verify against xAI docs before implementing |
| DeepSeek | [ASSUMED] — OpenAI-compat; DeepSeek-V3 supports JSON mode | Verify against DeepSeek docs |
| OpenRouter | Full — passes json_object through to underlying model |
[CITED: openrouter.ai/docs/guides/features/structured-outputs] — quality varies by underlying model |
| Ollama | Partial — parameter accepted but some models ignore it | [ASSUMED based on known behavior] — fallback parse_classification() is D-02 |
| LMStudio | Partial — same as Ollama (OpenAI compat, model-dependent) | [ASSUMED] — fallback parse_classification() is D-02 |
Gemini note: The OpenAI-compat endpoint does NOT support response_format={"type": "json_object"} (string mode). It supports the structured Pydantic/schema form via response_format=CalendarEvent (using client.beta.chat.completions.parse()). For GenericOpenAIProvider which uses completions.create() with response_format={"type":"json_object"}, Gemini will likely ignore the parameter and return free-form text. [CITED: ai.google.dev/gemini-api/docs/openai]. Implication for planning: GenericOpenAIProvider should have a supports_json_mode: bool = True class-level flag, and the Gemini preset should set it to False to fall back to parse_classification().
Common Pitfalls
Pitfall 1: OpenAI SDK 2.x Empty API Key Rejection
What goes wrong: AsyncOpenAI(api_key="") raises OpenAIError: Missing credentials in SDK 2.34.0+. Affects Ollama, LMStudio, and any GenericOpenAIProvider configured without a key.
Why it happens: SDK 2.34.0 added credential validation. Empty string is falsy.
How to avoid: Always pass a non-empty placeholder: api_key=api_key or "not-needed".
Warning signs: OpenAIError: Missing credentials during provider construction.
Pitfall 2: HKDF AlreadyFinalized if Instance is Reused
What goes wrong: Calling hkdf.derive() twice on the same HKDF object raises AlreadyFinalized.
Why it happens: The cryptography library marks HKDF instances as consumed after one derivation.
How to avoid: Create a fresh HKDF(...) instance on every _derive_ai_settings_key() call. Explicitly documented in cloud_utils.py — replicate the comment.
Warning signs: cryptography.exceptions.AlreadyFinalized in encryption tests.
Pitfall 3: Celery bind=True + asyncio.run() Retry Scope
What goes wrong: self.retry() must be called from the sync task function, not from inside asyncio.run(_run(...)). If you raise self.retry() inside _run(), it escapes the event loop cleanup and may corrupt async state.
Why it happens: asyncio.run() runs a full event loop; raising a Celery Retry exception mid-loop is not supported.
How to avoid: _run() raises a custom _ClassificationError; the outer sync extract_and_classify() catches it and calls self.retry().
Warning signs: RuntimeError: Event loop is closed after retry.
Pitfall 4: is_active Dual-Write Race Condition
What goes wrong: Two concurrent admin requests both set is_active=True on different providers, resulting in two active providers.
Why it happens: Read-check then write is a TOCTOU pattern.
How to avoid: Use a single UPDATE system_settings SET is_active = (provider_id = $target) — updates all rows atomically. No Python-level read needed.
Pitfall 5: Anthropic output_config Schema Complexity Limit
What goes wrong: Complex classification schemas (many optional properties, union types) may hit Anthropic's grammar compilation limits: 24 optional parameters, 16 union-type parameters, 180s timeout.
Why it happens: Constrained decoding compiles the schema to a grammar; complex schemas take longer.
How to avoid: Keep the schema simple — assigned_topics: array[string], new_topic_suggestions: array[string], reasoning: string. All required. No unions. This is already what parse_classification() expects.
Pitfall 6: AdminAiConfigTab.vue Is Already Per-User (Wrong Scope)
What goes wrong: The existing AdminAiConfigTab.vue shows a per-user AI provider table (calling adminListUsers() and adminUpdateAiConfig(userId, ...)). Phase 7's admin AI Providers section is about global system-level provider settings, not per-user overrides.
Why it happens: The existing component was built for ADMIN-05 (per-user assignment). Phase 7 adds a different concept: global system_settings.
How to avoid: Add a new "Providers" sub-section in the admin AI Config tab for system-level settings. Keep the per-user assignment table. Do not confuse the two.
Pitfall 7: MAX_AI_CHARS Defined in Three Places
What goes wrong: MAX_AI_CHARS = 8_000 is defined in openai_provider.py, anthropic_provider.py, and classifier.py. Removing it from one but not the others leads to inconsistent truncation.
How to avoid: A single grep-and-delete step. The planner must include this as an explicit task action.
Code Examples
Anthropic output_config Usage
# Source: platform.claude.com/docs/en/build-with-claude/structured-outputs
import anthropic
_CLASSIFICATION_SCHEMA = {
"type": "object",
"properties": {
"assigned_topics": {"type": "array", "items": {"type": "string"}},
"new_topic_suggestions": {"type": "array", "items": {"type": "string"}},
"reasoning": {"type": "string"},
},
"required": ["assigned_topics", "new_topic_suggestions"],
"additionalProperties": False,
}
class AnthropicProvider(AIProvider):
def __init__(self, api_key: str, model: str, context_chars: int):
self._model = model
self._context_chars = context_chars
self._client = anthropic.AsyncAnthropic(api_key=api_key) # singleton
async def classify(self, document_text, existing_topics, system_prompt):
truncated = self._truncate(document_text)
topics_str = ", ".join(existing_topics) if existing_topics else "(none yet)"
user_msg = f"Existing topics: [{topics_str}]\n\nDocument text:\n{truncated}"
response = await self._client.messages.create(
model=self._model,
max_tokens=1024,
system=system_prompt,
messages=[{"role": "user", "content": user_msg}],
output_config={
"format": {"type": "json_schema", "schema": _CLASSIFICATION_SCHEMA}
},
)
# Fallback: if stop_reason is not "end_turn", response.content[0].text may be
# empty or partial — parse_classification handles gracefully.
raw = response.content[0].text if response.content else ""
return parse_classification(raw)
Celery Exponential Backoff
# Source: docs.celeryq.dev/en/stable/userguide/tasks.html (bind + manual countdown)
class _ClassificationError(Exception):
"""Raised by _run() for retryable classification failures only."""
@celery_app.task(
name="tasks.document_tasks.extract_and_classify",
bind=True,
max_retries=3,
)
def extract_and_classify(self, document_id: str) -> dict:
"""D-09: exponential backoff — 30s, 90s, 270s."""
try:
return asyncio.run(_run(document_id))
except _ClassificationError as exc:
countdowns = [30, 90, 270]
countdown = countdowns[min(self.request.retries, 2)]
raise self.retry(exc=exc, countdown=countdown)
Smart Truncation (D-13)
def _truncate(self, text: str) -> str:
"""First 60% + last 40% of context window (D-13)."""
limit = self._context_chars
if len(text) <= limit:
return text
head_len = int(limit * 0.6)
tail_len = limit - head_len
return text[:head_len] + "\n[...truncated...]\n" + text[-tail_len:]
Alembic Migration 0005
# backend/alembic/versions/0005_system_settings.py
"""Add system_settings table for AI provider configuration."""
from alembic import op
import sqlalchemy as sa
revision = "0005"
down_revision = "0004"
def upgrade():
op.create_table(
"system_settings",
sa.Column("id", sa.dialects.postgresql.UUID(as_uuid=True), primary_key=True,
server_default=sa.text("gen_random_uuid()")),
sa.Column("provider_id", sa.String, nullable=False),
sa.Column("api_key_enc", sa.Text, nullable=True),
sa.Column("base_url", sa.Text, nullable=True),
sa.Column("model_name", sa.Text, nullable=False, server_default=""),
sa.Column("context_chars", sa.Integer, nullable=False, server_default="8000"),
sa.Column("is_active", sa.Boolean, nullable=False, server_default="false"),
sa.Column("created_at", sa.TIMESTAMP(timezone=True), nullable=False,
server_default=sa.text("now()")),
sa.Column("updated_at", sa.TIMESTAMP(timezone=True), nullable=False,
server_default=sa.text("now()")),
sa.UniqueConstraint("provider_id", name="uq_system_settings_provider_id"),
)
def downgrade():
op.drop_table("system_settings")
get_provider() Registry Pattern
# backend/ai/__init__.py — O(1) lookup, no if/elif chain
from ai.provider_config import ProviderConfig
from ai.openai_provider import OpenAIProvider
from ai.anthropic_provider import AnthropicProvider
from ai.generic_openai_provider import GenericOpenAIProvider
_REGISTRY: dict[str, type] = {
"openai": OpenAIProvider,
"anthropic": AnthropicProvider,
"gemini": GenericOpenAIProvider,
"groq": GenericOpenAIProvider,
"xai": GenericOpenAIProvider,
"deepseek": GenericOpenAIProvider,
"openrouter": GenericOpenAIProvider,
"mistral": GenericOpenAIProvider,
"ollama": GenericOpenAIProvider,
"lmstudio": GenericOpenAIProvider,
}
def get_provider(config: ProviderConfig) -> AIProvider:
cls = _REGISTRY.get(config.provider_id)
if cls is None:
raise ValueError(f"Unknown AI provider: {config.provider_id!r}")
return cls(
api_key=config.api_key or "not-needed",
model=config.model,
base_url=config.base_url,
context_chars=config.context_chars,
)
Note: AnthropicProvider does not take base_url — its constructor signature will differ; handle in the factory.
Runtime State Inventory
This section is not applicable — this is a greenfield addition and refactor phase, not a rename/migration phase. No stored data keys, live service configs, or OS-registered state change names.
The one potential runtime state concern is the existing ai_provider and ai_model columns on the users table. These per-user overrides remain in place (ADMIN-05 still uses them). The new system_settings table provides the global system default that is used when a user's per-user override is NULL. The per-user fields do not change.
State of the Art
| Old Approach | Current Approach | When Changed | Impact |
|---|---|---|---|
| Anthropic tool_use for structured output | output_config.format.type="json_schema" with constrained decoding |
~Nov 2025 (GA by June 2026) | No more JSON parsing fragility for Anthropic; model cannot produce invalid output |
Celery return {"status": "classification_failed"} on exception |
self.retry(countdown=N, max_retries=3) then status update on exhaustion |
This phase | Classification failures now auto-retry with backoff |
_client() method creating AsyncOpenAI per call |
self._client = AsyncOpenAI(...) in __init__ |
This phase | Full httpx connection pool reuse |
Flat if/elif provider registry |
Dict-based _REGISTRY + ProviderConfig typed model |
This phase | O(1) lookup, type safety, easy to add providers |
MAX_AI_CHARS = 8_000 global constant |
Per-provider context_chars in ProviderConfig |
This phase | Different providers support wildly different context sizes |
| Env var only AI config | DB-backed system_settings with admin UI |
This phase | Runtime reconfiguration without restarting containers |
Deprecated/outdated:
MAX_AI_CHARSconstant inopenai_provider.py,anthropic_provider.py,classifier.py— remove all three.- The
_client()method pattern in both providers — replace with__init__singleton. get_provider(settings: dict)raw dict interface — replace withget_provider(config: ProviderConfig).- Hardcoded
if active == "anthropic"chain in__init__.py— replace with registry.
Docker Compose D-14 Status
Already done. Both backend and celery-worker services already have:
extra_hosts:
- "host.docker.internal:host-gateway"
This was added in a prior phase. D-14 requires no docker-compose changes. The LMStudio fix is fully covered by making the base URL configurable via system_settings (D-15).
Validation Architecture
workflow.nyquist_validation = true in .planning/config.json — this section is required.
Test Framework
| Property | Value |
|---|---|
| Framework | pytest + pytest-asyncio |
| Config file | pytest.ini or pyproject.toml (existing) |
| Quick run command | pytest backend/tests/test_classifier.py backend/tests/test_ai_providers.py -x |
| Full suite command | pytest backend/tests/ -v |
Phase Requirements → Test Map
| Req ID | Behavior | Test Type | Automated Command | File Exists? |
|---|---|---|---|---|
| D-01 | GenericOpenAIProvider passes response_format={"type":"json_object"} |
unit (mock) | pytest backend/tests/test_ai_providers.py::test_generic_openai_json_mode -x |
No — Wave 0 |
| D-03 | AnthropicProvider passes output_config.format.type="json_schema" |
unit (mock) | pytest backend/tests/test_ai_providers.py::test_anthropic_structured_output -x |
No — Wave 0 |
| D-04 | load_provider_config() reads from system_settings table | integration (real DB) | INTEGRATION=1 pytest backend/tests/test_ai_config.py::test_load_provider_config -x |
No — Wave 0 |
| D-05 | system_settings API keys encrypt/decrypt correctly | unit | pytest backend/tests/test_ai_config.py::test_api_key_encrypt_decrypt -x |
No — Wave 0 |
| D-05 | Admin GET /api/admin/ai-config never returns api_key_enc | integration | pytest backend/tests/test_admin_ai_config.py::test_get_never_returns_key -x |
No — Wave 0 |
| D-06 | get_provider() accepts ProviderConfig, not raw dict | unit | pytest backend/tests/test_ai_providers.py::test_get_provider_typed -x |
No — Wave 0 |
| D-07 | OpenAIProvider._client is not recreated on second call | unit | pytest backend/tests/test_ai_providers.py::test_client_singleton -x |
No — Wave 0 |
| D-09 | extract_and_classify retries with 30s/90s/270s backoff | unit (mock) | pytest backend/tests/test_document_tasks.py::test_retry_backoff -x |
No — Wave 0 |
| D-10 | After 3 retries doc.status = "classification_failed" | unit (mock) | pytest backend/tests/test_document_tasks.py::test_exhaustion_sets_failed_status -x |
No — Wave 0 |
| D-12 | context_chars respected per provider | unit | pytest backend/tests/test_ai_providers.py::test_context_chars_truncation -x |
No — Wave 0 |
| D-13 | Truncation is 60% head + 40% tail | unit | pytest backend/tests/test_ai_providers.py::test_smart_truncation -x |
No — Wave 0 |
| D-11 | POST /api/documents/{id}/classify re-queues Celery | integration | pytest backend/tests/test_documents.py::test_reclassify_requeues_celery -x |
No — Wave 0 |
| existing | parse_classification/parse_suggestions robustness | unit | pytest backend/tests/test_classifier.py -x |
Yes |
Sampling Rate
- Per task commit:
pytest backend/tests/test_classifier.py backend/tests/test_ai_providers.py -x - Per wave merge:
pytest backend/tests/ -v - Phase gate: Full suite green before
/gsd:verify-work
Wave 0 Gaps
backend/tests/test_ai_providers.py— covers D-01, D-03, D-06, D-07, D-12, D-13backend/tests/test_ai_config.py— covers D-04, D-05 (encryption round-trip + DB read)backend/tests/test_admin_ai_config.py— covers admin endpoint security (never returns key)backend/tests/test_document_tasks.py(new tests in existing file) — covers D-09, D-10
Note: test_classifier.py already exists and covers parse_classification() robustness — keep as-is.
Security Domain
Applicable ASVS Categories
| ASVS Category | Applies | Standard Control |
|---|---|---|
| V2 Authentication | No | — |
| V3 Session Management | No | — |
| V4 Access Control | Yes | get_current_admin on all /api/admin/ai-config endpoints |
| V5 Input Validation | Yes | Pydantic ProviderConfig validates all admin inputs |
| V6 Cryptography | Yes | HKDF/AES-GCM (same pattern as cloud credentials) |
Known Threat Patterns
| Pattern | STRIDE | Standard Mitigation |
|---|---|---|
| Admin endpoint returns api_key_enc | Information Disclosure | _ai_config_to_dict() whitelist — never include api_key_enc field |
| Prompt injection via base_url or model_name | Tampering | Pydantic validation; base_url validated against URL format; model_name is opaque string passed to provider |
| HKDF key reuse across domains | Elevation of Privilege | info=b"ai-provider-settings" ≠ info=b"cloud-credentials" — domain separation |
| Non-admin user reads system AI config | Unauthorized Access | Admin-only endpoints (get_current_admin dep) |
| Classification of attacker-controlled documents | Tampering | Document content is user's own; provider calls are server-side; no SSRF from provider base_url (it's admin-configured, not user-supplied) |
Admin endpoint invariant: GET /api/admin/ai-config and PUT /api/admin/ai-config must never return api_key_enc. Follow the _user_to_dict() whitelist pattern from admin.py — create _ai_config_to_dict(row) that includes provider_id, base_url, model_name, context_chars, is_active, updated_at but excludes api_key_enc. The API key field in the admin UI is write-only (masked input; never pre-filled from the server).
Environment Availability
| Dependency | Required By | Available | Version | Fallback |
|---|---|---|---|---|
| PostgreSQL | system_settings table | Yes (Docker) | 17-alpine | — |
| Redis | Celery retry queue | Yes (Docker) | 7-alpine | — |
| Python openai SDK ≥1.30 | GenericOpenAIProvider | Yes (requirements.txt) | Latest: 2.40.0 | — |
| Python anthropic SDK ≥0.26 | AnthropicProvider | Yes (requirements.txt) | Latest: 0.105.2 | — |
| Python celery ≥5.5.0 | bind=True retry | Yes (requirements.txt) | Latest: 5.6.3 | — |
| docker-compose extra_hosts | LMStudio/Ollama host access | Already present in docker-compose.yml | — | — |
Missing dependencies with no fallback: None.
Assumptions Log
| # | Claim | Section | Risk if Wrong |
|---|---|---|---|
| A1 | xAI/Grok API supports response_format={"type":"json_object"} |
JSON Compatibility Matrix | Planner must add supports_json_mode=False preset for xAI if not supported; fallback to parse_classification() |
| A2 | DeepSeek supports response_format={"type":"json_object"} |
JSON Compatibility Matrix | Same as A1 — fallback is already designed in D-02 |
| A3 | Context char defaults in PROVIDER_DEFAULTS table | Pattern 2 | Admin will reconfigure; only startup defaults; no functional risk |
| A4 | Ollama/LMStudio: response_format param is accepted but may be ignored by some models |
JSON Compatibility Matrix | Already handled by D-02 fallback; no implementation risk |
| A5 | anthropic SDK version ≥0.95 is installed (output_config support added around v0.95) | D-03 section | If older version: output_config kwarg raises TypeError; bump requirement to >=0.95.0 |
Note on A5: The requirements.txt pins anthropic>=0.26. The output_config parameter was introduced significantly after that. The planner must update the minimum pin to anthropic>=0.95.0 (or preferably the current latest 0.105.2).
Open Questions (RESOLVED)
-
xAI/Grok and DeepSeek JSON mode verification
- What we know: Both claim OpenAI-compat; industry pattern suggests they support
json_object - What's unclear: Neither was directly verified against official docs in this research session
- Recommendation: Planner should include a Wave 0 task to spot-check against official xAI and DeepSeek API docs, or treat both as [ASSUMED] and add
supports_json_modeflag that defaults toTruebut can be overridden - RESOLVED: treated as [ASSUMED]; supports_json_mode flag is the mitigation — planner accepts the risk. xAI and DeepSeek are configured with supports_json_mode=True in Plan 02; if a future user reports JSON-shape failures, flipping the SUPPORTS_JSON_MODE entry to False routes that provider through parse_classification() (D-02 fallback path). No blocking action required for Phase 7 execution.
- What we know: Both claim OpenAI-compat; industry pattern suggests they support
-
anthropic SDK minimum version for output_config
- What we know:
output_configis confirmed in current SDK (0.105.2) with no beta headers needed - What's unclear: The exact version when
output_configwas introduced (the oldoutput_format/beta-header path still works) - Recommendation: Bump
requirements.txttoanthropic>=0.95.0as part of Wave 1; ifoutput_formatparam (old alias) is still working, the older SDK might work too — but it's safer to pin to a known good version - RESOLVED: treated as [ASSUMED]; supports_json_mode flag is the mitigation — planner accepts the risk. Plan 02 pins anthropic>=0.95.0 in requirements.txt; if a stricter floor is needed it can be raised in a follow-up phase. The fallback path (parse_classification) handles any output_config regression.
- What we know:
-
Gemini-compat: use json_object or fall back?
- What we know: The OpenAI-compat endpoint does NOT document
json_objectstring mode; it supports Pydantic schema viaclient.beta.chat.completions.parse() - What's unclear: Whether
json_objectis silently ignored or raises an error - Recommendation: Implement Gemini preset with
supports_json_mode=False;GenericOpenAIProvider.classify()skipsresponse_formatfor these presets and relies onparse_classification()fallback - RESOLVED: treated as [ASSUMED]; supports_json_mode flag is the mitigation — planner accepts the risk. Plan 02 SUPPORTS_JSON_MODE["gemini"] = False routes Gemini through the parse_classification() fallback path. Test test_gemini_fallback_to_parse_classification in Plan 02 Task 4 enforces this contract.
- What we know: The OpenAI-compat endpoint does NOT document
Sources
Primary (HIGH confidence)
- Anthropic Structured Outputs Docs —
output_config.format, supported models, schema limits, no-beta-header confirmation - Gemini OpenAI Compatibility Docs — what is/is not supported on the compat endpoint
- Celery 5 Tasks Documentation —
bind=True,self.retry(),countdown,max_retries - OpenAI Python SDK API Reference — AsyncOpenAI singleton/connection pool recommendation
Secondary (MEDIUM confidence)
- Groq Structured Outputs —
json_objectsupport confirmed - Mistral JSON Mode —
response_format={"type":"json_object"}confirmed - OpenRouter Structured Outputs —
json_objectpassthrough confirmed - openai-python Issue #3224 — empty api_key breaking change in 2.34.0
- openai-python Issue #1725 — connection pool reuse requires singleton client
- Anthropic SDK Client Architecture — AsyncAnthropic httpx lifecycle
Tertiary (LOW confidence)
- Ollama/LMStudio JSON mode behaviour — [ASSUMED] from known community patterns; not verified against official Ollama API docs in this session
- xAI/Grok json_object support — [ASSUMED]
- DeepSeek json_object support — [ASSUMED]
Metadata
Confidence breakdown:
- Standard stack: HIGH — all packages already in requirements.txt, versions verified on PyPI
- Anthropic structured output (D-03): HIGH — verified against official Anthropic docs
- Client lifecycle (D-07): HIGH — verified against openai-python issues + Anthropic SDK docs
- JSON mode matrix: MEDIUM — Groq/Mistral/OpenRouter verified; xAI/DeepSeek/Gemini partially assumed
- Celery retry pattern: HIGH — verified against official Celery 5 docs
- system_settings design: HIGH — based on established cloud_utils.py pattern in codebase
- Pitfalls: HIGH — openai SDK empty key bug directly cited from issue tracker
Research date: 2026-06-02 Valid until: 2026-09-01 (stable APIs; Anthropic structured output is GA, no longer beta)