mirror of
https://github.com/AlexsJones/llmfit.git
synced 2026-08-29 21:21:40 +00:00
Add ModelFormat enum (Gguf/Awq/Gptq/Mlx/Safetensors) and Vllm inference runtime. Pre-quantized models are detected from config.json by the scraper and filtered to CUDA/ROCm only (no Apple Silicon support yet). Dynamic re-quantization is skipped for fixed-precision formats. Co-Authored-By: Claude <noreply@anthropic.com> |
||
|---|---|---|
| .. | ||
| __pycache__ | ||
| install-openclaw-skill.sh | ||
| scrape_hf_models.py | ||
| test_api.py | ||
| update_models.sh | ||
| verify_models.py | ||