Slow Response Time for First Prompt When Using Certain Models (e.g., Llama3.1) #2377

Closed
opened 2025-11-11 15:05:58 -06:00 by GiteaMirror · 0 comments
Owner

Originally created by @fatmaT2001 on GitHub (Oct 14, 2024).

Description:
Hello OpenWebUI team,

I’ve noticed that certain models, such as Llama3.1, experience a significant delay in responding to the first prompt, particularly during the GPU connection phase. This issue doesn't seem to affect all models, as others connect and respond much faster during the initial interaction.

Details:

  • Models experiencing delay: Llama3.1 (and other similar models)
  • Issue: First prompt response time is considerably slower while connecting to the GPU, but subsequent prompts respond faster.
  • Behavior: The delay occurs only during the initial prompt, and the system performs normally afterward.

Question:

Could you please provide insight into why certain models, like Llama3.1, take longer to respond during the first prompt while connecting to the GPU, whereas other models don’t exhibit this behavior? Is there a difference in model initialization or other factors causing this delay?

Thank you for your assistance in helping me understand this behavior!

Originally created by @fatmaT2001 on GitHub (Oct 14, 2024). **Description**: Hello OpenWebUI team, I’ve noticed that certain models, such as **Llama3.1**, experience a significant delay in responding to the first prompt, particularly during the GPU connection phase. This issue doesn't seem to affect all models, as others connect and respond much faster during the initial interaction. ### Details: - **Models experiencing delay**: Llama3.1 (and other similar models) - **Issue**: First prompt response time is considerably slower while connecting to the GPU, but subsequent prompts respond faster. - **Behavior**: The delay occurs only during the initial prompt, and the system performs normally afterward. ### Question: Could you please provide insight into why certain models, like Llama3.1, take longer to respond during the first prompt while connecting to the GPU, whereas other models don’t exhibit this behavior? Is there a difference in model initialization or other factors causing this delay? Thank you for your assistance in helping me understand this behavior!
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: github-starred/open-webui#2377