I have searched for any existing and/or related issues.
I have searched for any existing and/or related discussions.
I am using the latest version of Open WebUI.
Installation Method
Docker
Open WebUI Version
v0.6.28
Ollama Version (if applicable)
No response
Operating System
Linux
Browser (if applicable)
No response
Confirmation
I have read and followed all instructions in README.md.
I am using the latest version of both Open WebUI and Ollama.
I have included the browser console logs.
I have included the Docker container logs.
I have provided every relevant configuration, setting, and environment variable used in my setup.
I have clearly listed every relevant configuration, custom setting, environment variable, and command-line option that influences my setup (such as Docker Compose overrides, .env values, browser settings, authentication configurations, etc).
I have documented step-by-step reproduction instructions that are precise, sequential, and leave nothing to interpretation. My steps:
Start with the initial platform/version/OS and dependencies used,
Specify exact install/launch/configure commands,
List URLs visited, user input (incl. example values/emails/passwords if needed),
Describe all options and toggles enabled or changed,
Include any files or environmental changes,
Identify the expected and actual result at each stage,
Ensure any reasonably skilled user can follow and hit the same issue.
Expected Behavior
When using GPT-OSS via llama.cpp, the thinking process of the models output should be encapsulated in OWUI "Thinking..." visual
Actual Behavior
Harmony format tokens such as "<|channel|>analysis<|message|>" and "|end|><|start|>assistant<|channel|>final<|message|>" are seen instead
Steps to Reproduce
Build llama.cpp 6464, launch llama-server as follows: ./llama-server -m /home/LLM/Models/ggml-org/gpt-oss-120b-GGUF/gpt-oss-120b-mxfp4-00001-of-00003.gguf -c 131072 -ngl 999 -b 2048 -ub 2048 -fa on --host 0.0.0.0 --port 8081 --chat-template-kwargs '{"reasoning_effort":"high"}', prompt "Tell me a random fun fact about the Roman Empire" (or any prompt), thinking component will not be parsed
Logs & Screenshots
Result with --jinja:
Result without --jinja:
--jinja is the recommended way to run GPT-OSS in llama.cpp as it allows passing --chat-template-kwargs '{"reasoning_effort":"high"}' to influence the reasoning effort, but including this flag results in no thinking tokens being shown in OWUI
Additional Information
No response
Originally created by @AbdullahMPrograms on GitHub (Sep 14, 2025).
Original GitHub issue: https://github.com/open-webui/open-webui/issues/17428
### Check Existing Issues
- [x] I have searched for any existing and/or related issues.
- [x] I have searched for any existing and/or related discussions.
- [x] I am using the latest version of Open WebUI.
### Installation Method
Docker
### Open WebUI Version
v0.6.28
### Ollama Version (if applicable)
_No response_
### Operating System
Linux
### Browser (if applicable)
_No response_
### Confirmation
- [x] I have read and followed all instructions in `README.md`.
- [x] I am using the latest version of **both** Open WebUI and Ollama.
- [x] I have included the browser console logs.
- [x] I have included the Docker container logs.
- [x] I have **provided every relevant configuration, setting, and environment variable used in my setup.**
- [x] I have clearly **listed every relevant configuration, custom setting, environment variable, and command-line option that influences my setup** (such as Docker Compose overrides, .env values, browser settings, authentication configurations, etc).
- [x] I have documented **step-by-step reproduction instructions that are precise, sequential, and leave nothing to interpretation**. My steps:
- Start with the initial platform/version/OS and dependencies used,
- Specify exact install/launch/configure commands,
- List URLs visited, user input (incl. example values/emails/passwords if needed),
- Describe all options and toggles enabled or changed,
- Include any files or environmental changes,
- Identify the expected and actual result at each stage,
- Ensure any reasonably skilled user can follow and hit the same issue.
### Expected Behavior
When using GPT-OSS via llama.cpp, the thinking process of the models output should be encapsulated in OWUI "Thinking..." visual
### Actual Behavior
Harmony format tokens such as "<|channel|>analysis<|message|>" and "|end|><|start|>assistant<|channel|>final<|message|>" are seen instead
### Steps to Reproduce
Build llama.cpp 6464, launch llama-server as follows: ./llama-server -m /home/LLM/Models/ggml-org/gpt-oss-120b-GGUF/gpt-oss-120b-mxfp4-00001-of-00003.gguf -c 131072 -ngl 999 -b 2048 -ub 2048 -fa on --host 0.0.0.0 --port 8081 --chat-template-kwargs '{"reasoning_effort":"high"}', prompt "Tell me a random fun fact about the Roman Empire" (or any prompt), thinking component will not be parsed
### Logs & Screenshots
Result with --jinja:
<img width="903" height="221" alt="Image" src="https://github.com/user-attachments/assets/48127cc0-2245-4d3e-89fc-da2db2331c26" />
Result without --jinja:
<img width="904" height="864" alt="Image" src="https://github.com/user-attachments/assets/cd347339-94b0-4ac2-86e3-f1d90ab62d33" />
--jinja is the recommended way to run GPT-OSS in llama.cpp as it allows passing --chat-template-kwargs '{"reasoning_effort":"high"}' to influence the reasoning effort, but including this flag results in no thinking tokens being shown in OWUI
### Additional Information
_No response_
GiteaMirror
added the bug label 2026-05-18 03:12:17 -05:00
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Originally created by @AbdullahMPrograms on GitHub (Sep 14, 2025).
Original GitHub issue: https://github.com/open-webui/open-webui/issues/17428
Check Existing Issues
Installation Method
Docker
Open WebUI Version
v0.6.28
Ollama Version (if applicable)
No response
Operating System
Linux
Browser (if applicable)
No response
Confirmation
README.md.Expected Behavior
When using GPT-OSS via llama.cpp, the thinking process of the models output should be encapsulated in OWUI "Thinking..." visual
Actual Behavior
Harmony format tokens such as "<|channel|>analysis<|message|>" and "|end|><|start|>assistant<|channel|>final<|message|>" are seen instead
Steps to Reproduce
Build llama.cpp 6464, launch llama-server as follows: ./llama-server -m /home/LLM/Models/ggml-org/gpt-oss-120b-GGUF/gpt-oss-120b-mxfp4-00001-of-00003.gguf -c 131072 -ngl 999 -b 2048 -ub 2048 -fa on --host 0.0.0.0 --port 8081 --chat-template-kwargs '{"reasoning_effort":"high"}', prompt "Tell me a random fun fact about the Roman Empire" (or any prompt), thinking component will not be parsed
Logs & Screenshots
Result with --jinja:

Result without --jinja:
--jinja is the recommended way to run GPT-OSS in llama.cpp as it allows passing --chat-template-kwargs '{"reasoning_effort":"high"}' to influence the reasoning effort, but including this flag results in no thinking tokens being shown in OWUI
Additional Information
No response
@tjbck commented on GitHub (Sep 15, 2025):
@AbdullahMPrograms does llama.cpp not support
reasoning_contentfield?