Originally created by @gibru on GitHub (Jun 6, 2024).
Bug Report
Text Generation Hangs
Bug Summary:
Using Open WebUI, when prompting a model (e.g. Llama 70B), text generation stops after a while and the stop button is also unresponsive. Context length both tested default and 8K. Reloading the page allows to chat again, but the unfinished answer isn't saved. Using ollama directly doesn't have this issue (screenshots provided for comparison).
Steps to Reproduce:
Prompt the model for longer answers.
Expected Behavior:
The LLM is supposed to complete the answer and/or the stop button should allow for ending the text generation.
Environment
Open WebUI Version:0.2.5
Ollama (if applicable):1.4.1
Operating System: Linux / Docker Installation
Logs and Screenshots
Screenshots attached
Open WebUI example:
ollama (CLI) example:
Originally created by @gibru on GitHub (Jun 6, 2024).
# Bug Report
Text Generation Hangs
**Bug Summary:**
Using Open WebUI, when prompting a model (e.g. Llama 70B), text generation stops after a while and the stop button is also unresponsive. Context length both tested default and 8K. Reloading the page allows to chat again, but the unfinished answer isn't saved. Using ollama directly doesn't have this issue (screenshots provided for comparison).
**Steps to Reproduce:**
Prompt the model for longer answers.
**Expected Behavior:**
The LLM is supposed to complete the answer and/or the stop button should allow for ending the text generation.
## Environment
- **Open WebUI Version:0.2.5**
- **Ollama (if applicable):1.4.1**
- **Operating System: Linux / Docker Installation**
## Logs and Screenshots
Screenshots attached
**Open WebUI example:**

**ollama (CLI) example:**

Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Originally created by @gibru on GitHub (Jun 6, 2024).
Bug Report
Text Generation Hangs
Bug Summary:
Using Open WebUI, when prompting a model (e.g. Llama 70B), text generation stops after a while and the stop button is also unresponsive. Context length both tested default and 8K. Reloading the page allows to chat again, but the unfinished answer isn't saved. Using ollama directly doesn't have this issue (screenshots provided for comparison).
Steps to Reproduce:
Prompt the model for longer answers.
Expected Behavior:
The LLM is supposed to complete the answer and/or the stop button should allow for ending the text generation.
Environment
Logs and Screenshots
Screenshots attached
Open WebUI example:

ollama (CLI) example:
