Is that with vanilla llama.cpp or with the third-party llama-swap manager? Last time I checked llama-swap was still the go-to solution, although I admit I haven't looked into it further.
The router mode and matrix routing in llama.cpp is still early days and it can't easily juggle multiple models as easily as llama-swap so there's still benefits if you're using 24-32GB cards that can run multiple models simultaneously.
The gulf is quickly shrinking though and it seems like the need for llama-swap will disappear soon.
If you revisit my comment and pay attention to the opener:
> but you might not be aware that llama-server can do multi-model for a while now
you will see that the sentence structure clearly implies both a change compared with a prior state and also lack of any third-party thing.
So the answer to the question has already been encoded as text available.
_
I can see the desire for explicit validation though. For that, I would propose a sentence structure like
> Oh cool! That means that llama-swap is now superseded/no longer needed?
That shows that you've read and understand the message, gives you the double-check and might on top spark a conversation about how these solutions compare.
Plus that if the guy you're commenting too has spoken nonsense, they need to backpedal.
btw, llama-swap provides a nice UI for monitoring performance and logs, and even the ability to stop an infinite session that consumes GPU resources (sometimes that happens).
Does the llama.cpp UI provide the same? If not, it is too early to say that llama-swap is “superseded/no longer needed.”