Commit graph

43 commits

Author SHA1 Message Date
JimmyZQX
f5bb0dad59
Data Filtering with Gemma (#396)
* Add code for data filtering llm judge

* Ignore log file created on root (mainly for synthetic_data_generation.log)

* Fix metadata API compatibility issues by commenting out metadata tags in LLM API calls

- Commented out metadata.tags parameters in all LLM API calls across the codebase
- This fixes compatibility issues with custom LLM providers that don't support metadata
- Affects shades generation, topics generation, wiki generation, bio QA, and question generation
- Preserves the original code structure for future re-enabling if needed

* feat: add data filtering pipeline with Ollama integration

- Add MergedDataJudge class for intelligent data filtering using Ollama Gemma
- Integrate automatic Ollama CLI installation into project setup process
- Add DATA_FILTERING step to training pipeline with concurrent processing
- Include testing for MergedDataJudge in its local main() function
- Add Ollama dependency to pyproject.toml

* feat: add automatic Ollama model cleanup after data filtering

* Add logging for outputting data filtering parameters

* fix: adjust error handling for MergedDataJudge:
- Keep original merged.json unchanged when any error occurs
- Exit filtering process immediately on errors instead of continuing with defaults
- Ensure training pipeline continues safely even if data filtering fails

* Add frontend for data filtering pipeline

* resolve data filtering quality_level error by commenting out problematic fields, change TrainProcessService back to original class definition

* fix: quote unquoted shade icons to prevent JSON parsing errors

* Fixed wiki_res.json missing due to no database connection at wiki/base.py module import

* Added scoring reasoning as part of the merged data

* fix: filter ANSI escape sequences from Ollama logs in data filtering step

* fix: Add data filtering steps to cloud training to resolve KeyError

- Added 'Data Filtering' step to cloud training progress holder
- Added data filtering step execution in cloud training service
- Added data filtering parameters to cloud training routes
- Updated frontend to send data filtering parameters
- Fixed missing except clause in cloud training service

This resolves the KeyError: 'data_filtering' when switching from cloud to local training.
2025-08-15 11:19:12 +08:00
doubleBlack2
d96eda2e0d
Fix/new data fix (#394)
* fix

* fix error

* fix stop error

* fix stop error

* fix stop error

* fix stop error

* fix stop error
2025-07-03 18:56:37 +08:00
wuyuxiangX
7becdc4f84
Feature/new data (#393)
* new data pipiline

* new data pipiline

* shades and topic

---------

Co-authored-by: yanmuyuan <2216646664@qq.com>
2025-07-01 15:39:04 +08:00
wuyuxiangX
ea86f26af3
Feature/cloud upgrade (#388)
* feat: Updated the cloud service model list and fixed the error message in the training process service

* cloud serive add

* add deployed_name

* add name,deploy model

* add model,name

* feat: Update cloud model deployment information, add deployment status and related parameters

* remove qwen2.5

* feat: Updated the model name to Qwen3, adjusted related configurations and default values

* refactor: simplify chat data processing and update model training configuration

* Qwen 3

* feat: Add cloud service status check function to optimize model status management

* chore: update API endpoints and file paths for local development environment

* feat: store complete training parameters in local storage

* Fix/cloud service stop (#384)

* fix cloud service stop

* add pending status

* feature: Added pause status polling function and updated cloud training stop logic

* Stop logic modification

* feature: Added cloud training pause status check function to optimize training process control

* fix stop

* feature: Add language parameter to support multi-language training configuration

* add Chinese language, fix top

---------

Co-authored-by: wyx-hhhh <1360479992@qq.com>

* feat: update originPrompt to simplify response structure and enhance clarity

* load error

---------

Co-authored-by: yanmuyuan <2216646664@qq.com>
Co-authored-by: doubleBlack2 <108928143+doubleBlack2@users.noreply.github.com>
2025-06-23 14:21:23 +08:00
doubleBlack2
a01aaa98dc
Fix/cloud service stop (#384)
* fix cloud service stop

* add pending status

* feature: Added pause status polling function and updated cloud training stop logic

* Stop logic modification

* feature: Added cloud training pause status check function to optimize training process control

* fix stop

* feature: Add language parameter to support multi-language training configuration

* add Chinese language, fix top

---------

Co-authored-by: wyx-hhhh <1360479992@qq.com>
2025-06-17 15:07:39 +08:00
doubleBlack2
f3e4d289e6
Feature/cloud service (#383)
* Enhance GGUF model handling with timestamps, metadata and memory training status

* Check if is_trained exists

* fix

* cloud service

* Change the data type of the is_trained field to boolean and update the related logic to reflect this change

* Change the data type of the is_trained field to boolean and update the related logic to reflect this change

* Add gguf path to json file

* Added model selection function, updated model list acquisition logic, and enhanced model information display

* Update the model service startup logic, add integrity check for the model path, and support obtaining the model path from different fields

* Service Change

* full cloud service

* feat: implement async cloud training process with job tracking and API key management

* Progress bar modification

* feat: Add Local and Cloud Training Configuration Components

- Introduced LocalTrainingConfig component for configuring local training parameters.
- Updated TrainingConfiguration component to include tabs for Local and Cloud training configurations.
- Added API functions for setting and getting cloud service API keys.
- Created useCloudProviderStore for managing cloud provider configurations.
- Enhanced event utility to include a new event for showing cloud provider modal.

* Refactor cloud provider and training configuration components

- Updated CloudProviderModal to handle cloud service API key management.
- Replaced API key handling with model configuration updates in CloudProviderModal.
- Enhanced CloudTrainingConfig to manage cloud models based on API key availability.
- Introduced new cloud service functions for listing available models and managing training jobs.
- Modified LocalTrainingConfig to ensure default model selection and synchronization.
- Updated TrainingConfiguration to manage model switching between local and cloud environments.
- Refactored useCloudProviderStore to integrate cloud service API key handling.
- Adjusted useTrainingStore to prioritize model name selection based on the active environment.

* Stream Output

* feat: Enhance training configuration and progress components

- Updated LocalTrainingConfig to improve default model handling and avoid unnecessary updates.
- Introduced LocalTrainingProgress component to manage local training progress display.
- Refactored TrainingConfiguration to support both local and cloud training types, including updated button text and actions.
- Modified TrainingProgress to conditionally render local or cloud training progress based on the selected training type.
- Added cloud service functions for starting training and managing job information.
- Adjusted training parameter interfaces to ensure consistency across local and cloud models.

* Stream response change

* feat: Enhance cloud training and inference capabilities

- Updated TrainingProgress component to handle cloud training progress data and job ID.
- Modified trainExposureModel to allow nullable path and added optional stageName.
- Enhanced useSSE hook to support cloud model inference with new parameters.
- Introduced CloudProgressData type to align cloud training progress with local training structure.
- Implemented cloud inference request handling with local knowledge retrieval in cloudService.
- Added utility functions for managing active cloud model state in cloudModelUtils.
- Updated cloud inference endpoint to support local knowledge retrieval before cloud inference.
- Refactored advanced chat service to utilize new message structure for cloud inference.
- Enhanced prompt strategies to incorporate knowledge retrieval based on user messages.

* feat: Delete the training parameter debugging information component

* Resume training at breakpoint

* Repair data redundancy

* Stop system modification

* fix error: reset training

* fix stop and reset

* Change chat reply format

* Enhance cloud and local service management with status tracking and improved progress reporting

- Implemented service status file management in cloud and local services to track active status and model information.
- Added endpoints to start and stop cloud services, including validation for existing services.
- Enhanced local service management with status checks and progress updates during document processing and chunk embedding.
- Introduced real-time progress tracking for document embedding and chunk processing, allowing for incremental updates.
- Improved error handling and logging throughout the service management processes.
- Refactored chat request handling to intelligently route between local and cloud services based on current status.

* feat:Cleaned up code comment

* translate Chinese comments to English in cloud service modules

* translate into chinese

* feat: Enhance cloud provider configuration and training management with API key handling and tab switching logic

* bug fix

* Add is_trained field modification in the cloud

* feat: Refactor training parameters management to separate local and cloud configurations

* feat: Update training parameter types to improve type safety and consistency

* feat: Add data synthesis mode to cloud training parameters and update related components

* feat: The document embedding part is restored to its original state

* refactor: optimize cloud training process with improved stop handling and file path updates

* feat: Update the default values and merging logic of cloud training parameters to ensure parameter consistency

* feat: Add API key preloading function to optimize the loading experience when the modal box is opened

* feat: Optimize CloudProviderModal component, add API key preloading and state management

* fix: Simplify cloud provider display by removing conditional rendering for Alibaba Cloud

* feat: Update .gitignore to include job_id.json and add .gitkeep for gguf directory

---------

Co-authored-by: wyx-hhhh <1360479992@qq.com>
2025-06-04 19:52:14 +08:00
CangWu
c89bae94b2
Feature/0509/add download current file (#368)
* feat:add current file to model download progress

* git ignore yarn

* feat: show current downloading

---------

Co-authored-by: Ye Xiangle <yexiangle@mail.mindverse.ai>
Co-authored-by: kevinaimonster <kevinaimonster@gmail.com>
2025-05-16 14:02:56 +08:00
wuyuxiangX
1f257184bd
Feature/fix duplicate files (#363)
* feat: file display and deletion related

* fix: file saving logic
2025-05-15 15:30:34 +08:00
KKKKKKKevin
5ceed311bf
Feature/add hint (#361)
* optimize doc

* better code

---------

Co-authored-by: kevinaimonster <kevinaimonster@gmail.com>
2025-05-14 20:26:45 +08:00
KKKKKKKevin
b4766a9d9d
UI Improvements and Thinking Model Configuration Enhancements (#353)
* Join AI Network -> Export your Second Me

* Default Synthesis Mode -> high
Default Epoch ->  3

* Set Thinking-Mode Default Value

* Better Display Of ReadMe

* default value of thinking mode

* Set Default value of enableL0Retrival to false
2025-05-13 19:42:34 +08:00
ryangyuan
c3855f37ad
Fix/0429/fix all log (#318)
* fix:fix return all log problem

* fix:delete no use code

* fix: add sse offset

* fix: change offset to string

* fix: fix more localStore

* fix: change log only

* fix: cancel offset

* fix: remove offset

* fix:delete no use code

* fix:add hertbeat

* fix: delete useless code

---------

Co-authored-by: Ye Xiangle <yexiangle@mail.mindverse.ai>
2025-05-07 15:53:25 +08:00
ryangyuan
b7f0cc7feb
Feat/0422/train l1 exposure (#319)
* feat:add get steps content(EXTRACT_DIMENSIONAL_TOPICS,MAP_ENTITY_NETWORK,DECODE_PREFERENCE_PATTERNS,AUGMENT_CONTENT_RETENTION)

* feat:add file_type

* feat: exponse train L1

* feat:jsonfy return data

* fix: jsonfy

* feat:add log

* fix: fix step change error

* feat:delete useless log

* fix:fix not import problem

* fix:fix old trainprocess init problem

* feat: Train Step Show Table

* feat:add L1_exposure_manager optimize code structure

* fix: fix bio return format & map_your_entity_network

* feat: show tip when resource empty

* add have_output & path

* fix: fix log problem

* feat: adjustment output ui

* fix: L1 exposure add loading

---------

Co-authored-by: Ye Xiangle <yexiangle@mail.mindverse.ai>
2025-05-06 16:02:22 +08:00
ryangyuan
5457a7a82a
fix: fix page overflow (#299)
* fix: add relative
2025-04-28 11:12:00 +08:00
ryangyuan
ef4c491d5f
Feat/0425/adjustment of training rule (#290)
* fix: adjustment status order

* fix: adjustment train status

* fix: split the status of service and train

* feat: adjustment train rule
2025-04-25 18:08:13 +08:00
ryangyuan
19adcac435
Feat/0423/train status (#287)
* fix: adjustment status order

* fix: adjustment train status

* fix: split the status of service and train
2025-04-25 17:46:37 +08:00
Zachary Pitroda
053090937d
Added CUDA support (#228)
* Add CUDA support

- CUDA detection
- Memory handling
- Ollama model release after training

* Fix logging issue

added cuda support flag so log accurately reflected cuda toggle

* Update llama.cpp rebuild

Changed llama.cpp to only check if cuda support is enabled and if so rebuild during the first build rather than each run

* Improved vram management

Enabled memory pinning and optimizer state offload

* Fix CUDA check

rewrote llama.cpp rebuild logic, added manual y/n toggle if user wants to enable cuda support

* Added fast restart and fixed CUDA check command

Added make docker-restart-backend-fast to restart the backend and reflect code changes without causing a full llama.cpp rebuild

Fixed make docker-check-cuda command to correctly reflect cuda support

* Added docker-compose.gpu.yml

Added docker-compose.gpu.yml to fix error on machines without nvidia gpu and made sure "\n" is added before .env modification

* Fixed cuda toggle

Last push accidentally broke cuda toggle

* Code review fixes

Fixed errors resulting from removed code:
- Added return save_path to end of save_hf_model function
- Rolled back download_file_with_progress function

* Update Makefile

Use cuda by default when using docker-restart-backend-fast

* Minor cleanup

Removed unnecessary makefile command and fixed gpu logging

* Delete .gpu_selected

* Simplified cuda training code

- Removed dtype setting to let torch automatically handle it
- Removed vram logging
- Removed Unnecessary/old comments

* Fixed gpu/cpu selection

Made "make docker-use-gpu/cpu" command work with .gpu_selected flag and changed "make docker-restart-backend-fast" command to respect flag instead of always using gpu

* Fix Ollama embedding error

Added custom exception class for Ollama embeddings, which seemed to be returning keyword arguments while the Python exception class only accepts positional ones

* Fixed model selection & memory error

Fixed training defaulting to 0.5B model regardless of selection and fixed "free(): double free detected in tcache 2" error caused by cuda flag being passed incorrectly
2025-04-25 10:20:36 +08:00
ryangyuan
f04916754c
feat: replace tutorial link (#268)
* feat: replace tutorial link

* replace video link

---------

Co-authored-by: kevin-mindverse <kevin@mindverse.ai>
2025-04-24 14:25:00 +08:00
ryangyuan
9fe511f0f2
Feature/0416/add thinking mode (#264)
* fix: modify thinking_model loading configuration

* feat: realize thinkModel ui

* feat:store

* feat: add combined_llm_config_dto

* add thinking_model_config & database migration

* directly add thinking model to user_llm_config

* delete thinking model repo dto service

* delete thinkingmodel table migration

* add is_cot config

* feat: allow define  is_cot

* feat: simplify logs info

* feat: add training model

* feat: fix is_cot problem

* fix: fix chat message

* fix: fix progress error

* fix: disable no settings thinking

* feat: add thinking warning

* fix: fix start service error

* feat:fix init trainparams problem

* feat: change playGround prompt

* feat: Add Dimension Mismatch Handling for ChromaDB (#157) (#207)

* Fix Issue #157

Add chroma_utils.py to manage chromaDB and added docs for explanation

* Add logging and debugging process

- Enhanced the`reinitialize_chroma_collections` function in`chroma_utils.py` to properly check if collections exist before attempting to delete them, preventing potential errors when collections don't exist.
- Improved error handling in the`_handle_dimension_mismatch` method in`embedding_service.py` by adding more robust exception handling and verification steps after reinitialization.
- Enhanced the collection initialization process in`embedding_service.py` to provide more detailed error messages and better handle cases where collections still have incorrect dimensions after reinitialization.
- Added additional verification steps to ensure that collection dimensions match the expected dimension after creation or retrieval.
- Improved logging throughout the code to provide more context in error messages, making debugging easier.

* Change topics_generator timeout to 30 (#263)

* quick fix

* fix: shade -> shade_merge_info (#265)

* fix: shade -> shade_merge_info

* add convert array

* quick fix import error

* add log

* add heartbeat

* new strategy

* sse version

* add heartbeat

* zh to en

* optimize code

* quick fix convert function

* Feat/new branch management (#267)

* feat: new branch management

* feat: fix multi-upload

* optimize contribute management

---------

Co-authored-by: Crabboss Mr <1123357821@qq.com>
Co-authored-by: Ye Xiangle <yexiangle@mail.mindverse.ai>
Co-authored-by: Xinghan Pan <sampan090611@gmail.com>
Co-authored-by: doubleBlack2 <108928143+doubleBlack2@users.noreply.github.com>
Co-authored-by: kevin-mindverse <kevin@mindverse.ai>
Co-authored-by: KKKKKKKevin <115385420+kevin-mindverse@users.noreply.github.com>
2025-04-24 14:19:23 +08:00
ryangyuan
fd64b4e5da
fix: fetch uploadInfo in homepage (#271) 2025-04-24 11:02:52 +08:00
ryangyuan
e1ae6f5039
fix: fix markdown (#253)
* fix: fix markdown

* feat:

---------

Co-authored-by: Llux <244421824@qq.com>
2025-04-17 16:41:52 +08:00
ryangyuan
177ac5d33a
fix: fix homePage fake loading (#248) 2025-04-16 19:39:01 +08:00
ryangyuan
7f82adc3fc
fix: change model name API (#246) 2025-04-16 17:04:57 +08:00
ryangyuan
2359630257
Feat/0414/adjustment train option (#236)
* feat: adjuesment train

* feat:optimize code structure

* feat:optimize params name

* fix: adjustment train log

* fix: fix train failed status

* fix: chosen model error

* feat: adjustment train

---------

Co-authored-by: Ye Xiangle <yexiangle@mail.mindverse.ai>
Co-authored-by: kevin-mindverse <kevin@mindverse.ai>
2025-04-16 15:50:06 +08:00
ryangyuan
26fd99fbf2
hotfix: apply scroll (#241)
* hotfix: apply scroll

* fix: adjustment padding
2025-04-16 12:01:56 +08:00
KKKKKKKevin
0e442ed17c
Fix/make meta data field optional (#237)
* fix:metadata to optional

* docker command check

* add exception stack
2025-04-15 19:51:33 +08:00
ryangyuan
e4eb21a173
Feature/0409/sidebar expand mcp (#233)
MCP

Co-authored-by: yanmuyuan <2216646664@qq.com>
2025-04-15 17:44:18 +08:00
ryangyuan
04a1785e40
feat: add Model Config Documentation (#223) 2025-04-15 10:27:49 +08:00
ryangyuan
000b39ac60
Fix: optimize progress order (#182)
* change progress order

* feat: change progress frontend type

* feat: update _load_progress for progress's new format

* fix: fix formatUnderscoreToName

* fix:simplify TrainProgress with direct JSON structure and mapping

* fix:add necessary accessors (getters and setters) to maintain compatibility with existing code while adopting the new data structure.

* fix: train log jump

* fix: fix current step update failed

* fix:fix stop problem

* fix trainprogress has no attribute status problem

* fix: merge master

---------

Co-authored-by: Ye Xiangle <yexiangle@mail.mindverse.ai>
2025-04-14 16:38:13 +08:00
yexiangle
4cbcc946aa
Adjustable Training Parameters & Checkpoint/Restore (#191)
* feat(trainprocess):add receive & get training params

* feat: support low/standard option for L2 data generate

* feat:improve training script invocation with direct bash command and internationalize comments

* fix: restore code

* feat: add error handling for GraphRAG index and modify model save frequency(epoch -> steps)

* feat: support low/medium/high mode for L2 data generation

* feat: set defaut value

* feat: add train params

* feat: use ScriptExecutor for training process and check return code

* fix: fix params round

* feat:use subprocess instead of script_executor to keep logs

* feat:hot fix name 'monitor_result' is not defined

* fix: delete useless message

* fix: fix params status

* feat: add status suspended

* feat:fix stop status

* fix: fix resume status

* fix: disable error

---------

Co-authored-by: Crabboss Mr <1123357821@qq.com>
Co-authored-by: ryangyuan <ryangyuan@mail.mindverse.ai>
2025-04-11 16:23:12 +08:00
CangWu
e3dc6da7de
fix: update next & postcss (#198) 2025-04-10 20:02:26 +08:00
CangWu
18654c5e6d
fix: fix some dependabot alerts (#196) 2025-04-10 19:33:40 +08:00
ryangyuan
b7184404f3
feat: add github stars (#192)
* feat: add github stars

* fix: fix code
2025-04-10 18:01:53 +08:00
ryangyuan
7f2cd1f199
feat: add chat endpoint tutorial (#184) 2025-04-10 11:16:48 +08:00
ryangyuan
e96fbc63e6
feat: add stopTrain loading (#136) 2025-04-01 20:32:47 +08:00
yuchengzhou
8741ef364e
Feature/0327/adapter chat to openai (#121)
* feat(chat): Chat Interface Protocol Modification to Align with OpenAI
2025-03-31 15:49:11 +08:00
KKKKKKKevin
f64ddc108b
feat: Add docker support (#99)
* feat:Add Docker Support

* translate
2025-03-30 15:15:32 +08:00
CangWu
391d488ee7
fix: copy to clipboard (#98) 2025-03-28 15:38:38 +08:00
Llux
2238b45c1d Merge branch 'master' of https://github.com/mindverse/Second-Me 2025-03-27 21:22:21 +08:00
Llux
42487d8bdc feat: config for host address 2025-03-27 21:22:14 +08:00
CangWu
8bb9ff11a0
Prevent WebGL runtime crash (#57)
* fix: enhance WebGL robustness

* del: delete README in frontend
2025-03-24 22:35:24 +08:00
ryangyuan
5f98ea2cd9 feat: add Network User Color 2025-03-21 11:29:17 +08:00
Llux
c6798e8ae8 chore: add overrides for braces 2025-03-20 14:06:51 +08:00
Kevin
7f7d64210e Initial commit 2025-03-20 00:37:54 +08:00