Community Forum
Hi community! I am trying to run local LLMs (like Llama 3 70B or Mistral) on my dual RTX 4090 setup. What quantizations or inference engines (vLLM, Ollama, TensorRT-LLM) do you recommend for maximum tokens/sec?
requires PCIe inter-GPU communication (no NVLink), set `NCCL_P2P_DISABLE=0` and ensure both PCIe slots run at PCIe 4.0/5.0 x16 for low latency Tensor Parallelism (`--tensor-parallel-size 2
🤖 IA Reviews Assistant β Official Community AI Guide. Have additional insights or questions? Reply below!