Note to first-time contributors: Please open a discussion post in Discussions to discuss your idea/fix with the community before creating a pull request.
Before submitting, make sure you've checked the following:
Target branch: This pull request targets the dev branch.
Description: A concise description of the changes is provided below.
Changelog: A README entry has been included.
Documentation: No user-facing behavior changes yet; documentation updates may follow depending on maintainers' feedback.
Dependencies: No new dependencies were introduced.
Testing: Manual testing performed to ensure the backend starts and builds successfully with the added utilities.
Agentic AI Code: This PR has been reviewed and manually validated before submission.
Code review: Self-review performed to ensure adherence to project conventions.
Design & Architecture: Changes are isolated utility additions and deployment scaffolding without altering existing runtime behavior.
Git Hygiene: This PR is atomic and focuses on deployment preparation utilities.
Title Prefix:chore
Description
This PR introduces backend utilities and deployment scaffolding to support running Open WebUI with an external HuggingFace-hosted MedGemma model server (e.g., served via vLLM on infrastructure such as RunPod).
The goal is to prepare the repository for integration with external LLM inference services rather than assuming local Ollama-based models.
These changes are non-breaking and primarily add optional utilities and configuration scaffolding that can be used in external deployments.
Key motivations:
enable Open WebUI deployments that rely on external model inference servers
prepare infrastructure for MedGemma-based medical document simplification workflows
introduce optional idle-shutdown support for GPU infrastructure (e.g., RunPod)
provide utilities for document chunking and prompt templating for long medical texts
None of these additions modify existing model providers or affect default Open WebUI behavior.
Changelog Entry
Description
Adds backend utilities and deployment scaffolding to support external model serving (e.g., MedGemma via vLLM) and infrastructure-aware deployments such as GPU spot instances.
Added
runpod_idle_shutdown.py utility to support optional automatic shutdown of idle GPU pods
document_chunking.py helper for splitting long medical documents into manageable chunks
medical document simplification prompt template
deployment scaffolding for running a MedGemma model server via vLLM
environment variable placeholders for external model server configuration
Changed
Extended .env.example with optional configuration variables for external model servers and GPU deployment environments.
Deprecated
None
Removed
None
Fixed
None
Security
No security-related changes.
Breaking Changes
None
Additional Information
These additions are intended for deployments where Open WebUI interacts with an external model inference server (e.g., HuggingFace models served through vLLM).
The utilities are optional and do not affect existing Open WebUI functionality.
Future work may include:
official configuration support for external HuggingFace model providers
deployment guides for GPU-hosted inference backends
Screenshots or Videos
N/A — backend utilities and deployment scaffolding only.
Contributor License Agreement
By submitting this pull request, I confirm that I have read and fully agree to the Contributor License Agreement (CLA), and I am providing my contributions under its terms.
🔄 This issue represents a GitHub Pull Request. It cannot be merged through Gitea due to API limitations.
## 📋 Pull Request Information
**Original PR:** https://github.com/open-webui/open-webui/pull/22371
**Author:** [@cadeferg](https://github.com/cadeferg)
**Created:** 3/7/2026
**Status:** ❌ Closed
**Base:** `main` ← **Head:** `medgemma-runpod-setup`
---
### 📝 Commits (3)
- [`6ef1770`](https://github.com/open-webui/open-webui/commit/6ef17701330b3889019ec419dbc313d361f5e951) Setting up custom docker image
- [`71564f2`](https://github.com/open-webui/open-webui/commit/71564f281a0777c8ca9320b268934832bab035c4) Update suggested prompts in config.py
- [`6fdab25`](https://github.com/open-webui/open-webui/commit/6fdab2512fedc8c9664b98ea1b8116aed8c1d0a1) prepare openwebui fork for medgemma runpod deployment
### 📊 Changes
**14 files changed** (+280 additions, -25 deletions)
<details>
<summary>View changed files</summary>
📝 `.env.example` (+27 -1)
📝 `Dockerfile` (+1 -0)
📝 `README.md` (+7 -0)
📝 `backend/open_webui/config.py` (+13 -13)
📝 `backend/open_webui/main.py` (+5 -0)
➕ `backend/open_webui/prompts/medical_simplification.txt` (+22 -0)
➕ `backend/open_webui/utils/document_chunking.py` (+21 -0)
➕ `backend/open_webui/utils/runpod_idle_shutdown.py` (+126 -0)
➕ `deploy/runpod-model/.env.example` (+4 -0)
➕ `deploy/runpod-model/README.md` (+11 -0)
➕ `deploy/runpod-model/run_vllm.sh` (+12 -0)
➕ `desktop.ini` (+2 -0)
➕ `docker-compose.medgemma-dev.yml` (+27 -0)
📝 `docker-compose.yaml` (+2 -11)
</details>
### 📄 Description
<!--
⚠️ CRITICAL CHECKS FOR CONTRIBUTORS (READ, DON'T DELETE) ⚠️
1. Target the `dev` branch. PRs targeting `main` will be automatically closed.
2. Do NOT delete the CLA section at the bottom. It is required for the bot to accept your PR.
-->
# Pull Request Checklist
### Note to first-time contributors: Please open a discussion post in Discussions to discuss your idea/fix with the community before creating a pull request.
**Before submitting, make sure you've checked the following:**
- [x] **Target branch:** This pull request targets the `dev` branch.
- [x] **Description:** A concise description of the changes is provided below.
- [x] **Changelog:** A README entry has been included.
- [ ] **Documentation:** No user-facing behavior changes yet; documentation updates may follow depending on maintainers' feedback.
- [x] **Dependencies:** No new dependencies were introduced.
- [x] **Testing:** Manual testing performed to ensure the backend starts and builds successfully with the added utilities.
- [x] **Agentic AI Code:** This PR has been reviewed and manually validated before submission.
- [x] **Code review:** Self-review performed to ensure adherence to project conventions.
- [x] **Design & Architecture:** Changes are isolated utility additions and deployment scaffolding without altering existing runtime behavior.
- [x] **Git Hygiene:** This PR is atomic and focuses on deployment preparation utilities.
- [x] **Title Prefix:** `chore`
---
# Description
This PR introduces backend utilities and deployment scaffolding to support running Open WebUI with an **external HuggingFace-hosted MedGemma model server** (e.g., served via vLLM on infrastructure such as RunPod).
The goal is to prepare the repository for integration with external LLM inference services rather than assuming local Ollama-based models.
These changes are **non-breaking** and primarily add optional utilities and configuration scaffolding that can be used in external deployments.
Key motivations:
- enable Open WebUI deployments that rely on **external model inference servers**
- prepare infrastructure for **MedGemma-based medical document simplification workflows**
- introduce **optional idle-shutdown support** for GPU infrastructure (e.g., RunPod)
- provide utilities for **document chunking and prompt templating** for long medical texts
None of these additions modify existing model providers or affect default Open WebUI behavior.
---
# Changelog Entry
### Description
Adds backend utilities and deployment scaffolding to support external model serving (e.g., MedGemma via vLLM) and infrastructure-aware deployments such as GPU spot instances.
### Added
- `runpod_idle_shutdown.py` utility to support optional automatic shutdown of idle GPU pods
- `document_chunking.py` helper for splitting long medical documents into manageable chunks
- medical document simplification prompt template
- deployment scaffolding for running a MedGemma model server via vLLM
- environment variable placeholders for external model server configuration
### Changed
- Extended `.env.example` with optional configuration variables for external model servers and GPU deployment environments.
### Deprecated
- None
### Removed
- None
### Fixed
- None
### Security
- No security-related changes.
### Breaking Changes
- **None**
---
### Additional Information
These additions are intended for deployments where Open WebUI interacts with an external model inference server (e.g., HuggingFace models served through vLLM).
The utilities are optional and do not affect existing Open WebUI functionality.
Future work may include:
- official configuration support for external HuggingFace model providers
- deployment guides for GPU-hosted inference backends
---
### Screenshots or Videos
N/A — backend utilities and deployment scaffolding only.
---
### Contributor License Agreement
By submitting this pull request, I confirm that I have read and fully agree to the [Contributor License Agreement (CLA)](https://github.com/open-webui/open-webui/blob/main/CONTRIBUTOR_LICENSE_AGREEMENT), and I am providing my contributions under its terms.
---
<sub>🔄 This issue represents a GitHub Pull Request. It cannot be merged through Gitea due to API limitations.</sub>
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
📋 Pull Request Information
Original PR: https://github.com/open-webui/open-webui/pull/22371
Author: @cadeferg
Created: 3/7/2026
Status: ❌ Closed
Base:
main← Head:medgemma-runpod-setup📝 Commits (3)
6ef1770Setting up custom docker image71564f2Update suggested prompts in config.py6fdab25prepare openwebui fork for medgemma runpod deployment📊 Changes
14 files changed (+280 additions, -25 deletions)
View changed files
📝
.env.example(+27 -1)📝
Dockerfile(+1 -0)📝
README.md(+7 -0)📝
backend/open_webui/config.py(+13 -13)📝
backend/open_webui/main.py(+5 -0)➕
backend/open_webui/prompts/medical_simplification.txt(+22 -0)➕
backend/open_webui/utils/document_chunking.py(+21 -0)➕
backend/open_webui/utils/runpod_idle_shutdown.py(+126 -0)➕
deploy/runpod-model/.env.example(+4 -0)➕
deploy/runpod-model/README.md(+11 -0)➕
deploy/runpod-model/run_vllm.sh(+12 -0)➕
desktop.ini(+2 -0)➕
docker-compose.medgemma-dev.yml(+27 -0)📝
docker-compose.yaml(+2 -11)📄 Description
Pull Request Checklist
Note to first-time contributors: Please open a discussion post in Discussions to discuss your idea/fix with the community before creating a pull request.
Before submitting, make sure you've checked the following:
devbranch.choreDescription
This PR introduces backend utilities and deployment scaffolding to support running Open WebUI with an external HuggingFace-hosted MedGemma model server (e.g., served via vLLM on infrastructure such as RunPod).
The goal is to prepare the repository for integration with external LLM inference services rather than assuming local Ollama-based models.
These changes are non-breaking and primarily add optional utilities and configuration scaffolding that can be used in external deployments.
Key motivations:
None of these additions modify existing model providers or affect default Open WebUI behavior.
Changelog Entry
Description
Adds backend utilities and deployment scaffolding to support external model serving (e.g., MedGemma via vLLM) and infrastructure-aware deployments such as GPU spot instances.
Added
runpod_idle_shutdown.pyutility to support optional automatic shutdown of idle GPU podsdocument_chunking.pyhelper for splitting long medical documents into manageable chunksChanged
.env.examplewith optional configuration variables for external model servers and GPU deployment environments.Deprecated
Removed
Fixed
Security
Breaking Changes
Additional Information
These additions are intended for deployments where Open WebUI interacts with an external model inference server (e.g., HuggingFace models served through vLLM).
The utilities are optional and do not affect existing Open WebUI functionality.
Future work may include:
Screenshots or Videos
N/A — backend utilities and deployment scaffolding only.
Contributor License Agreement
By submitting this pull request, I confirm that I have read and fully agree to the Contributor License Agreement (CLA), and I am providing my contributions under its terms.
🔄 This issue represents a GitHub Pull Request. It cannot be merged through Gitea due to API limitations.