I'm experiencing an issue with multiple Ollama API configurations across 3 servers running identical Ollama models. When testing with 2 different chats to observe query distribution across Ollama instances, I notice that instead of parallel processing, it completes one chat response before starting the second. Could anyone with similar setup share their OpenWebUI and Ollama configuration settings? Thanks for the help.
Originally created by @jotaperez3 on GitHub (Jan 26, 2025).
Original GitHub issue: https://github.com/open-webui/open-webui/issues/8930
I'm experiencing an issue with multiple Ollama API configurations across 3 servers running identical Ollama models. When testing with 2 different chats to observe query distribution across Ollama instances, I notice that instead of parallel processing, it completes one chat response before starting the second. Could anyone with similar setup share their OpenWebUI and Ollama configuration settings? Thanks for the help.

Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Originally created by @jotaperez3 on GitHub (Jan 26, 2025).
Original GitHub issue: https://github.com/open-webui/open-webui/issues/8930
I'm experiencing an issue with multiple Ollama API configurations across 3 servers running identical Ollama models. When testing with 2 different chats to observe query distribution across Ollama instances, I notice that instead of parallel processing, it completes one chat response before starting the second. Could anyone with similar setup share their OpenWebUI and Ollama configuration settings? Thanks for the help.