Sirius Claw
|
35dc7f01ac
|
fix: bump pinned vLLM version to v0.7.3 to support tool calling arguments
|
2026-04-20 06:58:13 +00:00 |
|
Sirius Claw
|
43eb51d5b7
|
feat: optimize cost via L4 spot instances and vLLM performance tuning
|
2026-04-20 05:38:36 +00:00 |
|
Sirius Claw
|
cecd14123e
|
chore: optimize vLLM parameters for A100 by removing artificial limits
|
2026-04-20 02:09:15 +00:00 |
|
sirius0xdev
|
46f576dcda
|
dial in vllm
|
2026-04-20 02:04:04 +00:00 |
|
sirius0xdev
|
0dbf223df0
|
fix vllm config
|
2026-04-20 01:49:36 +00:00 |
|
sirius0xdev
|
ee3cf7c7ee
|
dial in vllm settings
|
2026-04-19 23:41:15 +00:00 |
|
sirius0xdev
|
28e701e27f
|
add tools to vllm
|
2026-04-19 23:33:43 +00:00 |
|
sirius0xdev
|
d3d18ab28a
|
model switch
|
2026-04-18 22:29:44 +00:00 |
|
sirius0xdev
|
8aad1555a5
|
try new model
|
2026-04-18 22:25:26 +00:00 |
|
sirius0xdev
|
8ec068c841
|
vllm errors
|
2026-04-18 22:12:29 +00:00 |
|
sirius0xdev
|
3a924b4a10
|
fix vllm args
|
2026-04-18 22:08:56 +00:00 |
|
sirius0xdev
|
2fbd6bdf81
|
change resource limits
|
2026-04-18 22:04:41 +00:00 |
|
sirius0xdev
|
741a07bbc4
|
vllm errors models size might be too big
|
2026-04-18 21:58:39 +00:00 |
|
sirius0xdev
|
197212aacb
|
fix vllm errors
|
2026-04-18 21:53:00 +00:00 |
|
SiriusClaw
|
1ebab4f6e5
|
fix: enforce eager execution to stop gemma 2 CUDA graph crash
|
2026-04-18 21:32:03 +00:00 |
|
SiriusClaw
|
603465033b
|
feat: switch infrastructure to supergemma4-26b-abliterated-multimodal
|
2026-04-18 21:17:29 +00:00 |
|
SiriusClaw
|
1cc2d8e112
|
fix: remove XFORMERS attention backend override to fix Venice gibberish
|
2026-04-18 20:52:26 +00:00 |
|
sirius0xdev
|
73a4ee6408
|
Merge pull request #18 from sirius0xdev/fix/vllm-trust-remote-code-tokenizer
fix: revert tokenizer mode to auto
|
2026-04-18 16:35:51 -04:00 |
|
SiriusClaw
|
6df1e12ab0
|
fix: revert tokenizer mode to auto
|
2026-04-18 20:34:15 +00:00 |
|
sirius0xdev
|
518127c711
|
Merge pull request #17 from sirius0xdev/fix/vllm-mistral-regex
fix: enforce mistral tokenizer mode
|
2026-04-18 16:11:01 -04:00 |
|
SiriusClaw
|
0f5d0b17e5
|
fix: enforce mistral tokenizer mode to bypass regex corruption bug
|
2026-04-18 20:09:35 +00:00 |
|
SiriusClaw
|
9b7fc44ab2
|
feat: attach persistent volume claim to vLLM to cache HuggingFace weights
|
2026-04-18 20:07:30 +00:00 |
|
SiriusClaw
|
576ba503b0
|
fix: set vLLM dtype to auto to resolve weight casting corruption
|
2026-04-18 19:55:55 +00:00 |
|
SiriusClaw
|
7826ff0398
|
fix: adjust max-model-len and memory utilization for Venice
|
2026-04-18 19:35:20 +00:00 |
|
SiriusClaw
|
bd6416e6a6
|
fix: add trust-remote-code to vLLM args for mistral architecture
|
2026-04-18 19:01:53 +00:00 |
|
SiriusClaw
|
24a512e99d
|
fix: switch to Dolphin-Mistral-24B-Venice-Edition
|
2026-04-18 18:12:34 +00:00 |
|
SiriusClaw
|
f8bc8973a1
|
feat: upgrade brain to uncensored Dolphin Qwen2 72B AWQ
|
2026-04-18 17:57:32 +00:00 |
|
SiriusClaw
|
5560d2b4af
|
fix: switch to official Qwen2.5 32B Instruct model due to missing dolphin repo
|
2026-04-18 17:55:00 +00:00 |
|
SiriusClaw
|
1623ccb6e5
|
fix: correct huggingface model identifier for dolphin 32b and remove awq requirement for a100
|
2026-04-18 17:22:37 +00:00 |
|
SiriusClaw
|
b7b69195d0
|
refactor: migrate vLLM brain to A100 80GB node pool
|
2026-04-18 16:53:09 +00:00 |
|
SiriusClaw
|
cf909206f5
|
fix: correct kubernetes toleration effect syntax for vLLM deployment
|
2026-04-18 16:05:15 +00:00 |
|
SiriusClaw
|
55a3e9eb28
|
refactor: upgrade vLLM brain to uncensored Qwen 2.5 32B AWQ
|
2026-04-18 15:26:19 +00:00 |
|
SiriusClaw
|
e68774da48
|
refactor: switch vLLM deployment to uncensored Dolphin-Gemma model
|
2026-04-18 15:09:21 +00:00 |
|
SiriusClaw
|
3792be4262
|
feat: deploy vLLM Gemma model on RTX 6000 for SiriusClaw
|
2026-04-18 14:28:22 +00:00 |
|