open-notebook

mirror of https://github.com/lfnovo/open-notebook.git synced 2026-04-30 04:20:02 +00:00

Author	SHA1	Message	Date
Luis Novo	98eb6ed202	fix: use sync get_state() for SqliteSaver compatibility (#519 ) SqliteSaver does not support async methods like aget_state(). Use asyncio.to_thread() to run the sync get_state() call from async context, maintaining compatibility with the existing sync graph invocations. Closes #509	2026-01-31 19:25:11 -03:00
Luis Novo	5b2c97cab7	Fix re-embedding issues and improve retry strategy (#515 ) * fix: filter empty content in rebuild embeddings queries Update collect_items_for_rebuild() to properly filter out items with empty or whitespace-only content before submitting embedding jobs. Changes: - Sources: add string::trim(full_text) != '' filter - Notes: add string::trim(content) != '' filter - Insights: add content != none AND string::trim(content) != '' filter (previously had no content filter at all) This prevents unnecessary job submissions that would fail validation in the individual embed commands. Ref #513 * feat: add command_id to embedding error logs Add get_command_id() helper to extract command_id from execution context. Include command_id in error logs for all embedding commands: - embed_note_command - embed_insight_command - embed_source_command - create_insight_command This makes it easier to trace failed embedding jobs back to specific command records in the database. Ref #513 * fix: improve logging for embedding commands Log improvements: - Add command_id to all embedding error logs for traceability - Transaction conflicts in repo_insert now log at DEBUG (not ERROR) - Embedding API errors log at DEBUG, only ERROR when retries exhausted - Friendlier retry messages: "This will be retried automatically" - Include model name and command_id in generate_embeddings errors Files changed: - commands/embedding_commands.py: command_id in logs, friendlier messages - open_notebook/database/repository.py: DEBUG for transaction conflicts - open_notebook/utils/embedding.py: DEBUG logging, pass-through command_id Ref #513 * fix: correct field names in rebuild embeddings status endpoint The API status endpoint was looking for wrong field names: - sources_processed → sources_submitted - notes_processed → notes_submitted - insights_processed → insights_submitted - processed_items → jobs_submitted - failed_items → failed_submissions The command outputs "_submitted" because embedding happens async (we count jobs submitted, not items processed). Ref #513 * fix: update rebuild UI text to reflect async job submission Changed terminology from "Completed/processed" to "Jobs Submitted" since the rebuild command submits embedding jobs for async processing, not completing them synchronously. Updated in all locales: en-US, pt-BR, zh-CN, zh-TW, ja-JP Ref #513 * refactor: migrate retry strategy from allowlist to blocklist - Change from `retry_on: [RuntimeError, ...]` to `stop_on: [ValueError]` - This is more resilient: new exception types auto-retry by default - Simplified exception handling: ValueError = permanent, else = retry - Transient errors logged at DEBUG (surreal-commands logs final failure) - Permanent errors (ValueError) logged at ERROR Ref #513	2026-01-31 18:55:01 -03:00
Luis Novo	03f9edfec2	feat: use standard HTTP_PROXY/HTTPS_PROXY environment variables (#499 ) Update proxy configuration to use industry-standard environment variables (HTTP_PROXY, HTTPS_PROXY, NO_PROXY) instead of custom variables. The underlying libraries (esperanto, content-core, podcast-creator) now automatically detect proxy settings from these standard variables. - Bump content-core>=1.14.1 (fixes #494) - Bump esperanto>=2.18 - Bump podcast-creator>=0.9 - Update documentation with new proxy configuration	2026-01-29 23:31:02 -03:00
Fauzira Alpiandi	9adf70d18d	feat: message counting for chat sessions (#430 )	2026-01-29 23:00:22 -03:00
Luis Novo	4e411e0488	feat: add cascade deletion for notebooks with delete preview (#471 ) * feat: decrease chunking size for maximum ollama compatibility * docs: improve i18n info on Claude.md * feat: add cascade deletion for notebooks with delete preview - Add Notebook.get_delete_preview() to show counts of affected items - Add Notebook.delete(delete_exclusive_sources) for cascade deletion - Always delete notes when notebook is deleted - Allow user to choose: delete or keep exclusive sources - Shared sources are always unlinked but never deleted - Add NotebookDeleteDialog component with radio button options - Add delete-preview API endpoint - Update delete endpoint with delete_exclusive_sources param - Add i18n support for all 5 locales Closes #77 * docs: remove harcoded config settings	2026-01-25 14:56:14 -03:00
Luis Novo	d8006ff5cb	feat: content-type aware chunking and unified embedding (#444 ) * feat: content-type aware chunking and unified embedding - Add chunking.py with HTML, Markdown, and plain text detection - Add embedding.py with mean pooling for large content - Create dedicated commands: embed_note, embed_insight, embed_source - Use fire-and-forget pattern for embedding via submit_command() - Refactor rebuild_embeddings_command to delegate to individual commands - Remove legacy commands and needs_embedding() methods - Reduce chunk size to 1500 chars for Ollama compatibility - Update CLAUDE.md documentation for new architecture Fixes #350, #142 * fix: address code review issues - Note.save() now returns command_id for tracking embedding jobs - Add length check after generate_embeddings() to fail fast on mismatch - Add numpy as explicit dependency (was transitive) - Remove hardcoded chunk sizes from docstrings * docs: address code review comments - Rename "SYNC PATH" to "DOMAIN MODEL PATH" in embedding router - Add test_chunking.py and test_embedding.py to Testing Strategy - Clarify auto-embedding behavior for each domain model * fix: clean thinking tags from prompt graph output Adds clean_thinking_content() to prompt.py to handle extended thinking models that return <think>...</think> tags. This fixes empty titles when saving notes from chat. * chore: remove local docker-compose from git * fix(frontend): handle null parent_id in search results Add defensive check for null parent_id in search results to prevent "Cannot read properties of null (reading 'split')" error. This can happen with orphaned records in the database. * fix: cascade delete embeddings and insights when source is deleted When deleting a Source, now also deletes associated: - source_embedding records - source_insight records This prevents orphaned records that cause null parent_id errors in vector search results. * fix: add cleanup for orphan embedding/insight records in migration 10 Deletes source_embedding and source_insight records where the linked source no longer exists (source.id = NONE). * chore: bump esperanto to 2.16 Increases ctx_num for Ollama models to accommodate larger notebook context windows. See: https://github.com/lfnovo/esperanto/pull/69	2026-01-21 23:49:08 -03:00
MisonL	67dd85c928	Feat/localization tests docker (#371 ) * feat(i18n): complete 100% internationalization and fix Next.js 15 compatibility * feat(i18n): complete 100% internationalization coverage * chore(test): finalize component tests and project cleanup * test(logic): add unit tests for useModalManager hook * fix(test): resolve timeout in AppSidebar tests by mocking TooltipProvider * feat(i18n): comprehensive i18n audit, fixes for hardcoded strings, and complete zh-TW support * fix(i18n): resolve TypeScript warnings and improve translation hook stability - Remove unused useTranslation import from ConnectionGuard - Add ref-based checking state to prevent dependency cycles - Fix useTranslation hook to return empty string for undefined translations - Add comment for backward compatibility on ExtractedReference interface - Ensure .replace() string methods work safely with nested translation keys * feat(i18n): complete internationalization implementation with Docker deployment - Add LanguageLoadingOverlay component for smooth language transitions - Update all translation files (en-US, zh-CN, zh-TW) with improved terminology - Optimize Docker configuration for better performance - Update version check and config handling for i18n support - Fix route handling for language-specific content - Add comprehensive task documentation * fix(i18n): resolve localization errors, duplicates, and type issues * chore(i18n): finalize 100% internationalization coverage * chore(test): supplement i18n test cases and cleanup redundant files * fix(test): resolve lint type errors and finalize delivery documents * feat(i18n): finalize full internationalization and zh-TW localization * fix(frontend): add missing devDependency and fix build tsconfig * feat(ui): enhance sidebar hover effects with better visual feedback * fix(frontend): resolve accessibility, i18n, and lint issues - fix: add missing id, name, autocomplete attributes to dialog inputs - fix: add aria labels and DialogDescription for accessibility - fix: resolve uncontrolled component warning in SettingsForm - fix: correct duplicate 'Traditional Chinese' label in zh-TW locale - feat: add i18n support for podcast template names - chore: fix lint errors in Dialogs * fix: address all 21 PR feedback items from cubic-dev-ai bot Configuration: - Remove ignoreDuringBuilds flags from next.config.ts Testing: - Fix AppSidebar.test.tsx regex pattern and add missing assertion Logic: - Fix ConnectionGuard.tsx re-entry prevention logic Internationalization (I18n) - Translations: - Add missing keys: notebooks.archived, common.note/insight, accessibility keys - Add specific keys: sources.allSourcesDescShort, transformations.selectModel - Add singular/plural keys: podcasts.usedByCount_one/other, common.note/notes - Add common.created/updated with {time} placeholder Internationalization (I18n) - Usage: - SourcesPage: use allSourcesDescShort instead of string splitting - TransformationPlayground: use navigation.transformation and selectModel - CommandPalette: use dedicated keys instead of string concatenation - GeneratePodcastDialog: fix zh-TW date locale handling - NotebookHeader: correctly interpolate {time} placeholder - TransformationCard: use common.description instead of undefined key - ChatPanel/SpeakerProfilesPanel: implement proper pluralization - SystemInfo: correctly interpolate {version} placeholder - LanguageLoadingOverlay: use t.common.loading instead of hardcoded string - MessageActions: use specific error key cannotSaveNoteNoNotebook Other: - Fix SessionManager.tsx exhaustive-deps warning * fix: remove duplicate locale keys and add missing zh-CN translations - en-US: remove duplicate loading key (line 59) and addNew key (sources) - zh-CN: remove duplicate common keys (loading, note, insight, newSource, newNotebook, newPodcast) - zh-CN: remove duplicate accessibility.searchNotebooks key - zh-CN: remove duplicate sources.addNew key - zh-CN: remove duplicate navigation.transformation key - zh-CN: add missing usedByCount_one and usedByCount_other keys in podcasts - zh-TW: remove duplicate common keys (loading, note, insight, newSource, newNotebook, newPodcast) - zh-TW: remove duplicate accessibility.searchNotebooks key - zh-TW: remove duplicate sources.addNew key * docs: remove info.md * fix: remove duplicate notebook keys and unused ts-expect-error - zh-CN: remove duplicate notebooks keys (archived, archive, unarchive, deleteNotebook, deleteNotebookDesc) - zh-TW: remove duplicate notebooks keys (archived, archive, unarchive, deleteNotebook, deleteNotebookDesc) - GeneratePodcastDialog: remove unused @ts-expect-error directive * fix(a11y): fix unassociated labels in search page - Replace <Label> with role='group' + aria-labelledby for search type section - Replace <Label> with role='group' + aria-labelledby for search in section - Follows WAI-ARIA best practices for labeling form field groups * fix(a11y): fix unassociated labels across multiple components - search/page.tsx: use role='group' + aria-labelledby for search type and search in sections - RebuildEmbeddings.tsx: use role='group' + aria-labelledby for include checkboxes - TransformationPlayground.tsx: replace Label with span for non-form output label * chore: revert to npm stack and ensure i18n compatibility * chore: polish zh-TW translations for better idiomatic usage * fix: resolve linter errors (ruff import sort, mypy config duplicate) * style: apply ruff formatting * fix: finalize upstream compliance (Dockerfile.single, i18n hooks, docker-compose) * style: polish strings, fix timeout cleanup, and improve test mocks * fix: use relative imports in test setup to resolve IDE path errors * perf(docker): optimize build speed by removing apt-get upgrade and build tools - Remove apt-get upgrade from both builder and runtime stages (saves 10-15 min each) - Remove gcc/g++/make/git from builder (uv downloads pre-built wheels) - Add --no-install-recommends to minimize package footprint - Keep npm mirror (npmmirror.com) for faster frontend deps - Add npm registry config for reliable China network access Also includes: - fix(a11y): add missing labels and aria attributes to form fields - fix(i18n): add 2s safety timeout to LanguageLoadingOverlay - fix(i18n): add robustness checks to use-translation proxy Build time reduced from 2+ hours to ~34 minutes (~70% improvement) * fix(a11y): resolve 16 form field accessibility warnings in notebook and podcast pages * fix(a11y): resolve 4 button and 1 select field accessibility warnings in models page * fix(a11y): resolve redundant attributes and residual warnings in transformations and podcast forms * fix(i18n): deep fix for language switch hang using proxy protection and safer access * fix(a11y): add name attributes to ModelSelector, TransformationPlayground, and SourceDetailContent * fix: add missing Label import to SourceDetailContent * fix(i18n): use native react-i18next in LanguageLoadingOverlay to prevent hang during language switch * fix(i18n): rewrite use-translation Proxy with strict depth limit and expanded blocked props to prevent language switch hang * fix: add type assertion to fix TypeScript comparison error * fix(i18n): disable useSuspense to prevent thread hang during language resource loading * fix(i18n): add infinite loop detection circuit breaker to useTranslation hook * fix(i18n): update traditional chinese label to native script in en-US * feat: add new localization strings for notebook and note management. * fix: resolve config priority, docker build deps, and ui glitches * refactor: improve ui details and test coverage based on feedback * refactor: improve ui details (version check/lang toggle) and test coverage * fix: polish language matching and test cleanup * fix(test): update mocks to resolve timeouts and proxy errors * fix(frontend): restore tsconfig.json structure and enable IDE support for tests * fix: address PR review findings and resolve CI OIDC failure * fix: merge exception headers in custom handler * fix: comprehensive PR review remediations and async performance fixes * refactor: address all PR #371 review feedback - Docker: consolidate SURREAL_URL to docker.env, add single-container override - Security: restore apt-get upgrade in Dockerfile and Dockerfile.single - Create centralized getDateLocale helper (lib/utils/date-locale.ts) - Refactor 7 files to use getDateLocale helper - Revert config/route.ts to origin/main version - Move test files to co-located pattern (3 files) - Remove local useTranslation mock from ConfirmDialog.test.tsx - Simplify use-version-check to single useEffect pattern - Fix test import paths after moving to co-located pattern * fix: add jest-dom types for test files * fix: address remaining review issues - Add apt-get upgrade -y to Dockerfile.single backend-builder stage - Refactor ChatColumn.test.tsx: use 'as unknown as ReturnType<typeof hook>' instead of 'as any' - Use toBeInTheDocument() assertions instead of toBeDefined()	2026-01-15 13:51:05 -03:00
LUIS NOVO	71b8d13b24	docs: generate comprehensive CLAUDE.md reference documentation across codebase Create a hierarchical CLAUDE.md documentation system for the entire Open Notebook codebase with focus on concise, pattern-driven reference cards rather than comprehensive tutorials. ## Changes ### Core Documentation System - Updated `.claude/commands/build-claude-md.md` to distinguish between leaf and parent modules, with special handling for prompt/template modules - Established clear patterns: * Leaf modules (40-70 lines): Components, hooks, API clients * Parent modules (50-150 lines): Architecture, cross-layer patterns, data flows * Template modules: Pattern focus, not catalog listings ### Generated Documentation Created 15 CLAUDE.md reference files across the project: Frontend (React/Next.js) - frontend/src/CLAUDE.md: Architecture overview, data flow, three-tier design - frontend/src/lib/hooks/CLAUDE.md: React Query patterns, state management - frontend/src/lib/api/CLAUDE.md: Axios client, FormData handling, interceptors - frontend/src/lib/stores/CLAUDE.md: Zustand state persistence, auth patterns - frontend/src/components/ui/CLAUDE.md: Radix UI primitives, CVA styling Backend (Python/FastAPI) - open_notebook/CLAUDE.md: System architecture, layer interactions - open_notebook/ai/CLAUDE.md: Model provisioning, Esperanto integration - open_notebook/domain/CLAUDE.md: Data models, ObjectModel/RecordModel patterns - open_notebook/database/CLAUDE.md: Repository pattern, async migrations - open_notebook/graphs/CLAUDE.md: LangGraph workflows, async orchestration - open_notebook/utils/CLAUDE.md: Cross-cutting utilities, context building - open_notebook/podcasts/CLAUDE.md: Episode/speaker profiles, job tracking API & Other - api/CLAUDE.md: REST layer, service architecture - commands/CLAUDE.md: Async command handlers, job queue patterns - prompts/CLAUDE.md: Jinja2 templates, prompt engineering patterns (refactored) Project Root - CLAUDE.md: Project overview, three-tier architecture, tech stack, getting started ### Key Features - Zero duplication: Parent modules reference child CLAUDE.md files, don't repeat them - Pattern-focused: Emphasizes how components work together, not component catalogs - Scannable: Short bullets, code examples only when necessary (1-2 per file) - Practical: "How to extend" guides, quirks/gotchas for each module - Navigation: Root CLAUDE.md acts as hub pointing to specialized documentation ### Cleanup - Removed unused `batch_fix_services.py` - Removed deprecated `open_notebook/plugins/podcasts.py` - Updated .gitignore for documentation consistency ## Impact New contributors can now: 1. Read root CLAUDE.md for system architecture (5 min) 2. Jump to specific layer documentation (frontend, api, open_notebook) 3. Dive into module-specific patterns in child CLAUDE.md files (1 min per module) All documentation is lean, reference-focused, and avoids duplication.	2026-01-03 16:27:52 -03:00
Justin Florentine	869664a10b	fix: strip <think> tags from chat responses Add thinking content cleaning to notebook and source chat graphs. Previously, models that output <think>...</think> tags (like DeepSeek) or malformed variants without opening tags (like Nemotron) would leak reasoning content into user-visible responses. Changes: - chat.py: Clean AI response content before returning messages - source_chat.py: Same fix for source-specific chat - text_utils.py: Handle malformed output where opening <think> tag is missing but </think> is present 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2025-12-18 16:31:23 -05:00
Luis Novo	1a67f1f912	fix: enhance chat reference links and prevent text overflow (#173 ) This commit addresses two related issues in the chat interface: 1. Fix broken reference links (OSS-310) - Completely rewrote convertReferencesToMarkdownLinks() with greedy pattern matching - Now handles all edge cases: references after commas, nested brackets, bold markdown - Added visual icon indicators (FileText, Lightbulb, FileEdit) for reference types - Implemented proper error handling with toast notifications - Added validation for reference types and ID lengths 2. Fix long URL/text overflow (#172) - Added break-words and overflow-wrap classes to chat messages - Long URLs and text now wrap properly within chat bubbles - Applied fix consistently across source chat, notebook chat, and search results Technical Details: - Enhanced reference detection algorithm processes from end to start to preserve indices - Context analysis (50 chars before/after) determines original formatting - Icons are 12px, accessible, and themed appropriately - All changes pass linting and build successfully Files Modified: - frontend/src/lib/utils/source-references.tsx (core algorithm rewrite) - frontend/src/components/source/ChatPanel.tsx (error handling + text wrapping) - frontend/src/components/search/StreamingResponse.tsx (error handling + text wrapping) - open_notebook/utils/token_utils.py (ruff formatting fix) fixes #172	2025-10-19 15:38:59 -03:00
Luis Novo	aa593c60bd	feat: add persistent tiktoken cache to reduce re-downloads (#171 ) Some checks are pending Development Build / extract-version (push) Waiting to run Details Development Build / test-build-regular (push) Blocked by required conditions Details Development Build / test-build-single (push) Blocked by required conditions Details Development Build / summary (push) Blocked by required conditions Details Configure tiktoken to cache tokenizer encodings in ./data/tiktoken-cache instead of using system temp directory. This prevents re-downloading encoding files on every container restart and improves startup time. Changes: - Add TIKTOKEN_CACHE_DIR configuration in config.py - Set TIKTOKEN_CACHE_DIR environment variable in token_utils.py - Bump version to 1.0.7	2025-10-19 14:50:52 -03:00
LUIS NOVO	8b5daa86bc	fix: max tokens max is 8192 now	2025-10-18 13:21:53 -03:00
Luis Novo	b7e656a319	Version 1 (#160 ) New front-end Launch Chat API Manage Sources Enable re-embedding of all contents Sources can be added without a notebook now Improved settings Enable model selector on all chats Background processing for better experience Dark mode Improved Notes Improved Docs: - Remove all Streamlit references from documentation - Update deployment guides with React frontend setup - Fix Docker environment variables format (SURREAL_URL, SURREAL_PASSWORD) - Update docker image tag from :latest to :v1-latest - Change navigation references (Settings → Models to just Models) - Update development setup to include frontend npm commands - Add MIGRATION.md guide for users upgrading from Streamlit - Update quick-start guide with correct environment variables - Add port 5055 documentation for API access - Update project structure to reflect frontend/ directory - Remove outdated source-chat documentation files	2025-10-18 12:46:22 -03:00

13 commits