[GH-ISSUE #8930] Load Balancing Configuration #15330

Closed
opened 2026-04-19 21:34:37 -05:00 by GiteaMirror · 0 comments
Owner

Originally created by @jotaperez3 on GitHub (Jan 26, 2025).
Original GitHub issue: https://github.com/open-webui/open-webui/issues/8930

I'm experiencing an issue with multiple Ollama API configurations across 3 servers running identical Ollama models. When testing with 2 different chats to observe query distribution across Ollama instances, I notice that instead of parallel processing, it completes one chat response before starting the second. Could anyone with similar setup share their OpenWebUI and Ollama configuration settings? Thanks for the help.

Image

Originally created by @jotaperez3 on GitHub (Jan 26, 2025). Original GitHub issue: https://github.com/open-webui/open-webui/issues/8930 I'm experiencing an issue with multiple Ollama API configurations across 3 servers running identical Ollama models. When testing with 2 different chats to observe query distribution across Ollama instances, I notice that instead of parallel processing, it completes one chat response before starting the second. Could anyone with similar setup share their OpenWebUI and Ollama configuration settings? Thanks for the help. ![Image](https://github.com/user-attachments/assets/8237e4ad-b6f1-4355-8351-5da926f6383b)
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: github-starred/open-webui#15330