Home | Notifications | New Note | Local | Federated | Search | Logout
Note Detail
Ganbold@ganbold (2026-08-02 16:50:20)
honestly extremely annoyed with the current throughput per second for either prefill and decode. using cloud inference, it won't give me a time to breath. but using local inference, it gives me plenty of room to breath but I feel like my brain in stale state for a good few seconds which lead it to booting once the inference is done.
I am not sure what can I do with 4GB/6GB/8GB VRAM. kinda afraid with the model quality around that size even if it gives me speed. the amount of money to waste is just too scary
Reply