Commit graph

11 commits

Author SHA1 Message Date
wuyuxiangX
ea86f26af3
Feature/cloud upgrade (#388)
* feat: Updated the cloud service model list and fixed the error message in the training process service

* cloud serive add

* add deployed_name

* add name,deploy model

* add model,name

* feat: Update cloud model deployment information, add deployment status and related parameters

* remove qwen2.5

* feat: Updated the model name to Qwen3, adjusted related configurations and default values

* refactor: simplify chat data processing and update model training configuration

* Qwen 3

* feat: Add cloud service status check function to optimize model status management

* chore: update API endpoints and file paths for local development environment

* feat: store complete training parameters in local storage

* Fix/cloud service stop (#384)

* fix cloud service stop

* add pending status

* feature: Added pause status polling function and updated cloud training stop logic

* Stop logic modification

* feature: Added cloud training pause status check function to optimize training process control

* fix stop

* feature: Add language parameter to support multi-language training configuration

* add Chinese language, fix top

---------

Co-authored-by: wyx-hhhh <1360479992@qq.com>

* feat: update originPrompt to simplify response structure and enhance clarity

* load error

---------

Co-authored-by: yanmuyuan <2216646664@qq.com>
Co-authored-by: doubleBlack2 <108928143+doubleBlack2@users.noreply.github.com>
2025-06-23 14:21:23 +08:00
doubleBlack2
899e2cb752
Provide a selection of Chinese mirror websites for Huggingface (#360)
* Provide a selection of Chinese mirror websites for Huggingface

* Remove default values
2025-05-16 11:55:27 +08:00
Xiang Ying
4f6dffc64d
Update utils.py for progress (#365) 2025-05-15 15:27:29 +08:00
yingapple
da2b704ed4 fix(model tokenizer): just use the model tokenizer without anythink else. 2025-05-09 16:09:20 +08:00
Zachary Pitroda
053090937d
Added CUDA support (#228)
* Add CUDA support

- CUDA detection
- Memory handling
- Ollama model release after training

* Fix logging issue

added cuda support flag so log accurately reflected cuda toggle

* Update llama.cpp rebuild

Changed llama.cpp to only check if cuda support is enabled and if so rebuild during the first build rather than each run

* Improved vram management

Enabled memory pinning and optimizer state offload

* Fix CUDA check

rewrote llama.cpp rebuild logic, added manual y/n toggle if user wants to enable cuda support

* Added fast restart and fixed CUDA check command

Added make docker-restart-backend-fast to restart the backend and reflect code changes without causing a full llama.cpp rebuild

Fixed make docker-check-cuda command to correctly reflect cuda support

* Added docker-compose.gpu.yml

Added docker-compose.gpu.yml to fix error on machines without nvidia gpu and made sure "\n" is added before .env modification

* Fixed cuda toggle

Last push accidentally broke cuda toggle

* Code review fixes

Fixed errors resulting from removed code:
- Added return save_path to end of save_hf_model function
- Rolled back download_file_with_progress function

* Update Makefile

Use cuda by default when using docker-restart-backend-fast

* Minor cleanup

Removed unnecessary makefile command and fixed gpu logging

* Delete .gpu_selected

* Simplified cuda training code

- Removed dtype setting to let torch automatically handle it
- Removed vram logging
- Removed Unnecessary/old comments

* Fixed gpu/cpu selection

Made "make docker-use-gpu/cpu" command work with .gpu_selected flag and changed "make docker-restart-backend-fast" command to respect flag instead of always using gpu

* Fix Ollama embedding error

Added custom exception class for Ollama embeddings, which seemed to be returning keyword arguments while the Python exception class only accepts positional ones

* Fixed model selection & memory error

Fixed training defaulting to 0.5B model regardless of selection and fixed "free(): double free detected in tcache 2" error caused by cuda flag being passed incorrectly
2025-04-25 10:20:36 +08:00
wiley
7caa368a9a
feat:Add the functionality to download from ModelScope when downloading from Hugging Face fails. (#213) 2025-04-14 13:22:42 +08:00
justcrab
c39e662562
feat: support LongCoT mode for DeepSeek-R1 data synthesis and model training (#126)
* feat: support LongCoT mode for DeepSeek-R1 data synthesis and model training

* fix: restore .env and setting.yaml configuration files

* fix: restore .env from L2 configuration files

---------

Co-authored-by: Xiang Ying <yingxiang835@gmail.com>
2025-04-07 10:11:23 +08:00
GoForceX
9a914c1529 fix: Invalid progress file path on Windows (#145) 2025-04-03 10:35:01 +08:00
justcrab
e778ebf82f
feat(logging): separate training log (#83)
* desparate part of traing logs v1

* fix train.py log

* optimize logging manage && fix monitor log

* feat(logging): desparate part of training logs v2

* delete no use code

* merge conflic

* delete chinese

---------

Co-authored-by: Ye Xiangle <yexiangle@mail.mindverse.ai>
Co-authored-by: Crabboss Mr <crabbossmr@CrabbossdeMacBook-Air.local>
2025-03-27 10:12:11 +08:00
umutcrs
552ee09fce
security! (#62)
* Potential fix for code scanning alert no. 110: Uncontrolled data used in path expression

Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>

* Potential fix for code scanning alert no. 109: Uncontrolled data used in path expression

Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>

---------

Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>
2025-03-26 10:13:34 +08:00
Kevin
7f7d64210e Initial commit 2025-03-20 00:37:54 +08:00