mirror of
https://github.com/lfnovo/open-notebook.git
synced 2026-08-14 19:43:51 +00:00
Some checks are pending
Development Build / extract-version (push) Waiting to run
Development Build / changes (push) Waiting to run
Tests / Frontend Lint (push) Waiting to run
Tests / Backend Tests (push) Waiting to run
Tests / Backend Lint (push) Waiting to run
Tests / Backend Typecheck (push) Waiting to run
Development Build / build-regular (push) Blocked by required conditions
Development Build / build-single (push) Blocked by required conditions
Development Build / summary (push) Blocked by required conditions
Tests / Frontend Tests (push) Waiting to run
Tests / Frontend Build (push) Waiting to run
* fix(sources): fall back to auto when a selected engine's runtime is absent The content-processing engine choice is persisted in the database; the runtime that serves it (Docling, local Crawl4AI) is installed on demand from environment flags evaluated at boot. The two therefore drift: a redeploy that drops OPEN_NOTEBOOK_ENABLE_CRAWL4AI/_DOCLING, a volume moved to a new deployment, or a failed on-demand install all leave a stored selection pointing at a runtime that is not there. The source graph passed that selection straight to content-core, so every affected extraction failed with "Could not extract any text content from this source" - no mention of the engine, the runtime, or the flag that would fix it. For a URL engine set to crawl4ai this breaks URL ingestion entirely. The graph now checks runtime availability before honoring the stored engine and degrades to content-core's "auto" chain, logging a WARNING that names the engine and the env var that would enable it. Engines with no opt-in runtime (auto/simple/firecrawl/jina) are passed through untouched. The availability probes moved from api/routers/capabilities.py to open_notebook/utils/runtime_capabilities.py so the graph can use them without importing from the API layer; the capabilities endpoint keeps identical behavior and its tests follow the probes to their new home. Found by the smoke-e2e agent during v1.14.0 release testing, on a dev environment that was in exactly this state. Pre-existing since v1.13.0 (#1122 made the runtimes opt-in, #432 made the stored selection take effect), not a v1.14.0 regression. * docs(changelog): record the unavailable-engine fallback fix
40 lines
1.4 KiB
Python
40 lines
1.4 KiB
Python
"""
|
|
Capabilities Router
|
|
|
|
Reports the runtime availability of the opt-in heavy extraction engines
|
|
(Docling, Crawl4AI local) so the frontend can gate the corresponding engine
|
|
options and the OCR toggle. These runtimes are installed on demand at container
|
|
startup (see scripts/docker-entrypoint.sh and
|
|
docs/7-DEVELOPMENT/decisions/ADR-007-optin-runtimes.md), so this endpoint probes
|
|
what is *actually* importable/reachable rather than trusting the enable flags.
|
|
|
|
The probes themselves live in open_notebook.utils.runtime_capabilities because
|
|
the source-processing graph needs the same signal (it must not pass content-core
|
|
an engine whose runtime is absent).
|
|
|
|
Endpoints:
|
|
- GET /capabilities - Availability of Docling and Crawl4AI runtimes
|
|
"""
|
|
|
|
from fastapi import APIRouter
|
|
|
|
from api.models import CapabilitiesResponse
|
|
from open_notebook.utils.runtime_capabilities import (
|
|
crawl4ai_local_ready,
|
|
crawl4ai_remote_configured,
|
|
docling_available,
|
|
)
|
|
|
|
router = APIRouter(prefix="/capabilities", tags=["capabilities"])
|
|
|
|
|
|
@router.get("", response_model=CapabilitiesResponse)
|
|
async def get_capabilities():
|
|
"""Report which opt-in extraction runtimes are available in this container."""
|
|
crawl4ai_remote = crawl4ai_remote_configured()
|
|
crawl4ai_local = crawl4ai_local_ready()
|
|
return CapabilitiesResponse(
|
|
docling_available=docling_available(),
|
|
crawl4ai_available=crawl4ai_local or crawl4ai_remote,
|
|
crawl4ai_remote_configured=crawl4ai_remote,
|
|
)
|