[GH-ISSUE #5064] ollama-cuda: Incoherent responses using llama2-7b-chat/llama3:instruct after first rest request #3198

Open
opened 2026-04-12 13:41:30 -05:00 by GiteaMirror · 0 comments
Owner

Originally created by @adrianmarinopeya on GitHub (Jun 15, 2024).
Original GitHub issue: https://github.com/ollama/ollama/issues/5064

What is the issue?

I am using Manjaro Linux. A few days ago, I received a new update to ollama-cuda-0.1.41-1-x86_64.pkg.tar.zst version.
I noticed that when making a request to ollama API using llama2-7b-chat or llama3:instruct models, the initial
response is valid with a response time of 8 seconds. However, when I repeat the request multiple times, the
responses start to degrade to the point of becoming incoherent. However, the response time is significantly shorter that first time.

I solve this issue making a downgrade to ollama-cuda-0.1.38-1-x86_6.pkg.tar.zst. The responses became coherent again.

OS

Linux

GPU

Nvidia

CPU

AMD

Ollama version

0.1.41

Originally created by @adrianmarinopeya on GitHub (Jun 15, 2024). Original GitHub issue: https://github.com/ollama/ollama/issues/5064 ### What is the issue? I am using Manjaro Linux. A few days ago, I received a new update to `ollama-cuda-0.1.41-1-x86_64.pkg.tar.zst` version. I noticed that when making a request to ollama API using `llama2-7b-chat` or `llama3:instruct` models, the initial response is valid with a response time of 8 seconds. However, when I repeat the request multiple times, the responses start to degrade to the point of becoming incoherent. However, the response time is significantly shorter that first time. I solve this issue making a downgrade to `ollama-cuda-0.1.38-1-x86_6.pkg.tar.zst`. The responses became coherent again. ### OS Linux ### GPU Nvidia ### CPU AMD ### Ollama version 0.1.41
GiteaMirror added the bug label 2026-04-12 13:41:30 -05:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: github-starred/ollama#3198