Community Forum
quick question guys before i upgrade my GPU... will a RTX 4060 Ti 16GB run DeepSeek R1 14B Q4 smoothly with 8k context or should i save up for a used 3090 24GB?
16GB will fit 14B Q4 easily with around 45 tokens/sec on llama.cpp. but if u wanna use 16k context or 32B quant then definitely get the used 3090 24gb, VRAM is king for LLMs.
thanks Mark! yeah I might just look for a good 3090 deal to be future proof for 32b models.