Note to first-time contributors: Please open a discussion post in Discussions and describe your changes before submitting a pull request.
Before submitting, make sure you've checked the following:
Target branch: Please verify that the pull request targets the dev branch.
Description: Provide a concise description of the changes made in this pull request.
Changelog: Ensure a changelog entry following the format of Keep a Changelog is added at the bottom of the PR description.
Documentation: Have you updated relevant documentation Open WebUI Docs, or other documentation sources?
Dependencies: Are there any new dependencies? Have you updated the dependency versions in the documentation?
Testing: Have you written and run sufficient tests to validate the changes?
Code review: Have you performed a self-review of your code, addressing any coding standard issues and ensuring adherence to the project's coding standards?
Prefix: To clearly categorize this pull request, prefix the pull request title using one of the following:
BREAKING CHANGE: Significant changes that may affect compatibility
build: Changes that affect the build system or external dependencies
ci: Changes to our continuous integration processes or workflows
chore: Refactor, cleanup, or other non-functional code changes
docs: Documentation update or addition
feat: Introduces a new feature or enhancement to the codebase
fix: Bug fix or error correction
i18n: Internationalization or localization changes
perf: Performance improvement
refactor: Code restructuring for better maintainability, readability, or scalability
style: Changes that do not affect the meaning of the code (white space, formatting, missing semi-colons, etc.)
test: Adding missing tests or correcting existing tests
WIP: Work in progress, a temporary label for incomplete or ongoing work
Changelog Entry
Description
Hi, after 09874ab83d, the FireCrawlLoader's source URLs are now correctly handled!
However, there’s still a small point to address. I strongly recommend changing FireCrawlLoader's default mode to 'scrape'. The 'scrape' mode focuses on fetching a single URL, which is consistent with the behavior of other web loaders. In contrast, selecting 'crawl' mode—particularly when using a self-hosted Firecrawl instance—can result in Firecrawl recursively fetching all linked pages from the given URL, leading to unacceptably long processing times for just one webpage.
backend/open_webui/retrieval/web/utils.py: SafeFireCrawlLoader.__init__ set default mode to scrape
Contributor License Agreement
By submitting this pull request, I confirm that I have read and fully agree to the CONTRIBUTOR_LICENSE_AGREEMENT, and I am providing my contributions under its terms.
🔄 This issue represents a GitHub Pull Request. It cannot be merged through Gitea due to API limitations.
## 📋 Pull Request Information
**Original PR:** https://github.com/open-webui/open-webui/pull/13177
**Author:** [@tth37](https://github.com/tth37)
**Created:** 4/23/2025
**Status:** ✅ Merged
**Merged:** 4/29/2025
**Merged by:** [@tjbck](https://github.com/tjbck)
**Base:** `dev` ← **Head:** `fix_firecrawl_loader_default_mode`
---
### 📝 Commits (1)
- [`8f7195c`](https://github.com/open-webui/open-webui/commit/8f7195cedaf84efb683866a6fae764cfc897d6d3) fix: FireCrawlLoader default mode to scrape
### 📊 Changes
**1 file changed** (+1 additions, -1 deletions)
<details>
<summary>View changed files</summary>
📝 `backend/open_webui/retrieval/web/utils.py` (+1 -1)
</details>
### 📄 Description
# Pull Request Checklist
### Note to first-time contributors: Please open a discussion post in [Discussions](https://github.com/open-webui/open-webui/discussions) and describe your changes before submitting a pull request.
**Before submitting, make sure you've checked the following:**
- [x] **Target branch:** Please verify that the pull request targets the `dev` branch.
- [x] **Description:** Provide a concise description of the changes made in this pull request.
- [x] **Changelog:** Ensure a changelog entry following the format of [Keep a Changelog](https://keepachangelog.com/) is added at the bottom of the PR description.
- [ ] **Documentation:** Have you updated relevant documentation [Open WebUI Docs](https://github.com/open-webui/docs), or other documentation sources?
- [ ] **Dependencies:** Are there any new dependencies? Have you updated the dependency versions in the documentation?
- [x] **Testing:** Have you written and run sufficient tests to validate the changes?
- [x] **Code review:** Have you performed a self-review of your code, addressing any coding standard issues and ensuring adherence to the project's coding standards?
- [x] **Prefix:** To clearly categorize this pull request, prefix the pull request title using one of the following:
- **BREAKING CHANGE**: Significant changes that may affect compatibility
- **build**: Changes that affect the build system or external dependencies
- **ci**: Changes to our continuous integration processes or workflows
- **chore**: Refactor, cleanup, or other non-functional code changes
- **docs**: Documentation update or addition
- **feat**: Introduces a new feature or enhancement to the codebase
- **fix**: Bug fix or error correction
- **i18n**: Internationalization or localization changes
- **perf**: Performance improvement
- **refactor**: Code restructuring for better maintainability, readability, or scalability
- **style**: Changes that do not affect the meaning of the code (white space, formatting, missing semi-colons, etc.)
- **test**: Adding missing tests or correcting existing tests
- **WIP**: Work in progress, a temporary label for incomplete or ongoing work
# Changelog Entry
### Description
Hi, after 09874ab83dcf7c39babde9e5142b4caf9b2c9193, the FireCrawlLoader's source URLs are now correctly handled!
However, there’s still a small point to address. I strongly recommend changing FireCrawlLoader's default mode to `'scrape'`. The `'scrape'` mode focuses on fetching a single URL, which is consistent with the behavior of other web loaders. In contrast, selecting `'crawl'` mode—particularly when using a **self-hosted** Firecrawl instance—can result in Firecrawl recursively fetching all linked pages from the given URL, leading to unacceptably long processing times for just one webpage.
Update: same concern was mentioned in #13169
### Changed
- `backend/open_webui/retrieval/web/utils.py`: `SafeFireCrawlLoader.__init__` set default mode to `scrape`
---
### Contributor License Agreement
By submitting this pull request, I confirm that I have read and fully agree to the [CONTRIBUTOR_LICENSE_AGREEMENT](CONTRIBUTOR_LICENSE_AGREEMENT), and I am providing my contributions under its terms.
---
<sub>🔄 This issue represents a GitHub Pull Request. It cannot be merged through Gitea due to API limitations.</sub>
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
📋 Pull Request Information
Original PR: https://github.com/open-webui/open-webui/pull/13177
Author: @tth37
Created: 4/23/2025
Status: ✅ Merged
Merged: 4/29/2025
Merged by: @tjbck
Base:
dev← Head:fix_firecrawl_loader_default_mode📝 Commits (1)
8f7195cfix: FireCrawlLoader default mode to scrape📊 Changes
1 file changed (+1 additions, -1 deletions)
View changed files
📝
backend/open_webui/retrieval/web/utils.py(+1 -1)📄 Description
Pull Request Checklist
Note to first-time contributors: Please open a discussion post in Discussions and describe your changes before submitting a pull request.
Before submitting, make sure you've checked the following:
devbranch.Changelog Entry
Description
Hi, after
09874ab83d, the FireCrawlLoader's source URLs are now correctly handled!However, there’s still a small point to address. I strongly recommend changing FireCrawlLoader's default mode to
'scrape'. The'scrape'mode focuses on fetching a single URL, which is consistent with the behavior of other web loaders. In contrast, selecting'crawl'mode—particularly when using a self-hosted Firecrawl instance—can result in Firecrawl recursively fetching all linked pages from the given URL, leading to unacceptably long processing times for just one webpage.Update: same concern was mentioned in #13169
Changed
backend/open_webui/retrieval/web/utils.py:SafeFireCrawlLoader.__init__set default mode toscrapeContributor License Agreement
By submitting this pull request, I confirm that I have read and fully agree to the CONTRIBUTOR_LICENSE_AGREEMENT, and I am providing my contributions under its terms.
🔄 This issue represents a GitHub Pull Request. It cannot be merged through Gitea due to API limitations.