[GH-ISSUE #7209] feat: disable RAG for model #30186

Closed
opened 2026-04-25 04:28:40 -05:00 by GiteaMirror · 12 comments
Owner

Originally created by @juancarlosm on GitHub (Nov 22, 2024).
Original GitHub issue: https://github.com/open-webui/open-webui/issues/7209

Originally assigned to: @tjbck on GitHub.

Feature Request

Problem Description:
When integrating Open-webui with agents or assistants (such as Assistants Code Interpreter), we only need access to the uploaded file, not the content or RAG processing. I managed how to access the uploaded file via open-webui API, and then I can use the file in my agents, but the unnecessary RAG processing, especially for larger files, consumes time and resources. Additionally, each new query for the same chat triggers again RAG for embeddings and vector search, which is inefficient if not needed.

Proposed Solution:
Introduce a parameter in the model configuration to disable RAG processing for specific models. This would streamline operations by bypassing unnecessary file processing.

Alternatives Considered:
I attempted using functions with self.file_handler = True, but this approach did not prevent RAG processing.

Related Issues:
Open-webui Issue #3556

Originally created by @juancarlosm on GitHub (Nov 22, 2024). Original GitHub issue: https://github.com/open-webui/open-webui/issues/7209 Originally assigned to: @tjbck on GitHub. # Feature Request **Problem Description:** When integrating Open-webui with agents or assistants (such as [Assistants Code Interpreter](https://platform.openai.com/docs/assistants/tools/code-interpreter)), we only need access to the uploaded file, not the content or RAG processing. I managed how to access the uploaded file via open-webui API, and then I can use the file in my agents, but the unnecessary RAG processing, especially for larger files, consumes time and resources. Additionally, each new query for the same chat triggers again RAG for embeddings and vector search, which is inefficient if not needed. **Proposed Solution:** Introduce a parameter in the model configuration to disable RAG processing for specific models. This would streamline operations by bypassing unnecessary file processing. **Alternatives Considered:** I attempted using functions with `self.file_handler = True`, but this approach did not prevent RAG processing. **Related Issues:** [Open-webui Issue #3556](https://github.com/open-webui/open-webui/issues/3556)
Author
Owner

@oatmealm commented on GitHub (Nov 22, 2024):

Similarly, and I'm not sure if it belongs here, when I've instructed a model to randomly pick facts from the RAG, I don't need it to run a similarity query, rather pass the documents and let the model do the work it needs to.

I've been able to get some models to follow such a prompt: select a fact from the documents and ask me a question, then correct my reply and provide more information.

I think right now I'll need to write a pipeline for this to work as expected, since tften the content the rag pipelines returns is not complete and the model can't find all the information it needs. It might be able to make a question but not find the answer in the sources.

<!-- gh-comment-id:2493597966 --> @oatmealm commented on GitHub (Nov 22, 2024): Similarly, and I'm not sure if it belongs here, when I've instructed a model to randomly pick facts from the RAG, I don't need it to run a similarity query, rather pass the documents and let the model do the work it needs to. I've been able to get some models to follow such a prompt: select a fact from the documents and ask me a question, then correct my reply and provide more information. I think right now I'll need to write a pipeline for this to work as expected, since tften the content the rag pipelines returns is not complete and the model can't find all the information it needs. It might be able to make a question but not find the answer in the sources.
Author
Owner

@CWrecker commented on GitHub (Jan 10, 2025):

It would also be nice to turn this and vector storage off during file upload for unsupported files (e.g. MP4).

#6246

<!-- gh-comment-id:2584082533 --> @CWrecker commented on GitHub (Jan 10, 2025): It would also be nice to turn this and vector storage off during file upload for unsupported files (e.g. MP4). #6246
Author
Owner

@rhajou commented on GitHub (Feb 12, 2025):

@juancarlosm did you find any solution around that? Facing the same issue.
In addition, one quick question, how did you manage to get the file path / url using your custom Agent?

<!-- gh-comment-id:2652932823 --> @rhajou commented on GitHub (Feb 12, 2025): @juancarlosm did you find any solution around that? Facing the same issue. In addition, one quick question, how did you manage to get the file path / url using your custom Agent?
Author
Owner

@vinismarques commented on GitHub (Feb 13, 2025):

I am also very interested in having the option not to use RAG. Models can handle everything in the context window most of the time, especially the new Google models.

<!-- gh-comment-id:2656704595 --> @vinismarques commented on GitHub (Feb 13, 2025): I am also very interested in having the option not to use RAG. Models can handle everything in the context window most of the time, especially the new Google models.
Author
Owner

@EmilianoGarciaLopez commented on GitHub (Feb 20, 2025):

Also very interested

<!-- gh-comment-id:2670338687 --> @EmilianoGarciaLopez commented on GitHub (Feb 20, 2025): Also very interested
Author
Owner

@wangjiyang commented on GitHub (Feb 26, 2025):

Need this feature also, I added a pipeline to process files, however this RAG feature is always trying to save vector information to vector db, which may takes very long time and failed for http request timeout. This feature is not always required since I need to process it by pipelines.

<!-- gh-comment-id:2685693721 --> @wangjiyang commented on GitHub (Feb 26, 2025): Need this feature also, I added a pipeline to process files, however this RAG feature is always trying to save vector information to vector db, which may takes very long time and failed for http request timeout. This feature is not always required since I need to process it by pipelines.
Author
Owner

@tjbck commented on GitHub (Feb 28, 2025):

You can set BYPASS_EMBEDDING_AND_RETRIEVAL to false from the admin settings.

<!-- gh-comment-id:2690257029 --> @tjbck commented on GitHub (Feb 28, 2025): You can set `BYPASS_EMBEDDING_AND_RETRIEVAL` to false from the admin settings.
Author
Owner

@mkagit commented on GitHub (Mar 21, 2025):

You can set BYPASS_EMBEDDING_AND_RETRIEVAL to false from the admin settings.

It still appends the RAG task instruction to the system prompt even when BYPASS_EMBEDDING_AND_RETRIEVAL is false

<!-- gh-comment-id:2743589887 --> @mkagit commented on GitHub (Mar 21, 2025): > You can set `BYPASS_EMBEDDING_AND_RETRIEVAL` to false from the admin settings. It still appends the RAG task instruction to the system prompt even when BYPASS_EMBEDDING_AND_RETRIEVAL is false
Author
Owner

@mkagit commented on GitHub (Mar 21, 2025):

I want to send the attached file directly to the endpoint. Also when attaching an audio file it's triggers the transcribe engine. I want to send it to the endpoint.

<!-- gh-comment-id:2743598569 --> @mkagit commented on GitHub (Mar 21, 2025): I want to send the attached file directly to the endpoint. Also when attaching an audio file it's triggers the transcribe engine. I want to send it to the endpoint.
Author
Owner

@rgaricano commented on GitHub (Mar 21, 2025):

Image

<!-- gh-comment-id:2743760924 --> @rgaricano commented on GitHub (Mar 21, 2025): ![Image](https://github.com/user-attachments/assets/a0b1b3fe-5024-411c-8e7e-6079cb3d7646)
Author
Owner

@ColbyB722 commented on GitHub (Mar 24, 2025):

It still appends the RAG task instruction to the system prompt even when BYPASS_EMBEDDING_AND_RETRIEVAL is false

This is true. You can even see the RAG task instructions in the thinking tags of reasoning models like QwQ even when BYPASS_EMBEDDING_AND_RETRIEVAL is false

<!-- gh-comment-id:2748722677 --> @ColbyB722 commented on GitHub (Mar 24, 2025): > It still appends the RAG task instruction to the system prompt even when BYPASS_EMBEDDING_AND_RETRIEVAL is false This is true. You can even see the RAG task instructions in the thinking tags of reasoning models like QwQ even when `BYPASS_EMBEDDING_AND_RETRIEVAL` is false
Author
Owner

@marius10p commented on GitHub (Aug 6, 2025):

I believe this behavior is still true, even when BYPASS_EMBEDDING_AND_RETRIEVAL is false. The attached document still gets processed by the RAG task, even if the output of the task is not used, which takes too long on every prompt in a chat context.

<!-- gh-comment-id:3161682760 --> @marius10p commented on GitHub (Aug 6, 2025): I believe this behavior is still true, even when BYPASS_EMBEDDING_AND_RETRIEVAL is false. The attached document still gets processed by the RAG task, even if the output of the task is not used, which takes too long on every prompt in a chat context.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: github-starred/open-webui#30186