Commit graph

14 commits

Author SHA1 Message Date
SiriusClaw
6df1e12ab0 fix: revert tokenizer mode to auto 2026-04-18 20:34:15 +00:00
SiriusClaw
0f5d0b17e5 fix: enforce mistral tokenizer mode to bypass regex corruption bug 2026-04-18 20:09:35 +00:00
SiriusClaw
576ba503b0 fix: set vLLM dtype to auto to resolve weight casting corruption 2026-04-18 19:55:55 +00:00
SiriusClaw
7826ff0398 fix: adjust max-model-len and memory utilization for Venice 2026-04-18 19:35:20 +00:00
SiriusClaw
bd6416e6a6 fix: add trust-remote-code to vLLM args for mistral architecture 2026-04-18 19:01:53 +00:00
SiriusClaw
24a512e99d fix: switch to Dolphin-Mistral-24B-Venice-Edition 2026-04-18 18:12:34 +00:00
SiriusClaw
f8bc8973a1 feat: upgrade brain to uncensored Dolphin Qwen2 72B AWQ 2026-04-18 17:57:32 +00:00
SiriusClaw
5560d2b4af fix: switch to official Qwen2.5 32B Instruct model due to missing dolphin repo 2026-04-18 17:55:00 +00:00
SiriusClaw
1623ccb6e5 fix: correct huggingface model identifier for dolphin 32b and remove awq requirement for a100 2026-04-18 17:22:37 +00:00
SiriusClaw
b7b69195d0 refactor: migrate vLLM brain to A100 80GB node pool 2026-04-18 16:53:09 +00:00
SiriusClaw
cf909206f5 fix: correct kubernetes toleration effect syntax for vLLM deployment 2026-04-18 16:05:15 +00:00
SiriusClaw
55a3e9eb28 refactor: upgrade vLLM brain to uncensored Qwen 2.5 32B AWQ 2026-04-18 15:26:19 +00:00
SiriusClaw
e68774da48 refactor: switch vLLM deployment to uncensored Dolphin-Gemma model 2026-04-18 15:09:21 +00:00
SiriusClaw
3792be4262 feat: deploy vLLM Gemma model on RTX 6000 for SiriusClaw 2026-04-18 14:28:22 +00:00