208
Ollama is fine when it is just me - at what point does moving to vLLM actually pay for the setup pain?
Single 4090, currently running a 14b through Ollama behind a small FastAPI wrapper for my own use. I now need to open it to about eight colleagues for an internal tool, and in testing with four simultaneous requests it…