defsearch_google_pse(api_key:str,search_engine_id:str,query:str,count:int,filter_list:Optional[list[str]]=None,)->list[SearchResult]:"""Search using Google's Programmable Search Engine API and return the results as a list of SearchResult objects.
Handles pagination for counts greater than 10.
Args:
api_key (str): A Programmable Search Engine API key
search_engine_id (str): A Programmable Search Engine ID
query (str): The query to search for
count (int): The number of results to return (max 100, as PSE max results per query is 10 and max page is 10)
filter_list (Optional[list[str]], optional): A list of keywords to filter out from results. Defaults to None.
Returns:
list[SearchResult]: A list of SearchResult objects.
"""url="https://www.googleapis.com/customsearch/v1"headers={"Content-Type":"application/json"}all_results=[]start_index=1# Google PSE start parameter is 1-basedwhilecount>0:num_results_this_page=min(count,10)# Google PSE max results per page is 10params={"cx":search_engine_id,"q":query,"key":api_key,"num":num_results_this_page,"start":start_index,}response=requests.request("GET",url,headers=headers,params=params)response.raise_for_status()json_response=response.json()results=json_response.get("items",[])ifresults:# check if results are returned. If not, no more pages to fetch.all_results.extend(results)count-=len(results)# Decrement count by the number of results fetched in this page.start_index+=10# Increment start index for the next pageelse:break# No more results from Google PSE, break the loopiffilter_list:all_results=get_filtered_results(all_results,filter_list)return[SearchResult(link=result["link"],title=result.get("title"),snippet=result.get("snippet"),)forresultinall_results]
Originally created by @kulukami on GitHub (Feb 16, 2025).
# Feature Request
Now websearching is using ```requests``` raw requests. And it is quite useful if adding httpProxy option.
https://github.com/open-webui/open-webui/blob/main/backend/open_webui/retrieval/web/google_pse.py#L46
```python
def search_google_pse(
api_key: str,
search_engine_id: str,
query: str,
count: int,
filter_list: Optional[list[str]] = None,
) -> list[SearchResult]:
"""Search using Google's Programmable Search Engine API and return the results as a list of SearchResult objects.
Handles pagination for counts greater than 10.
Args:
api_key (str): A Programmable Search Engine API key
search_engine_id (str): A Programmable Search Engine ID
query (str): The query to search for
count (int): The number of results to return (max 100, as PSE max results per query is 10 and max page is 10)
filter_list (Optional[list[str]], optional): A list of keywords to filter out from results. Defaults to None.
Returns:
list[SearchResult]: A list of SearchResult objects.
"""
url = "https://www.googleapis.com/customsearch/v1"
headers = {"Content-Type": "application/json"}
all_results = []
start_index = 1 # Google PSE start parameter is 1-based
while count > 0:
num_results_this_page = min(count, 10) # Google PSE max results per page is 10
params = {
"cx": search_engine_id,
"q": query,
"key": api_key,
"num": num_results_this_page,
"start": start_index,
}
response = requests.request("GET", url, headers=headers, params=params)
response.raise_for_status()
json_response = response.json()
results = json_response.get("items", [])
if results: # check if results are returned. If not, no more pages to fetch.
all_results.extend(results)
count -= len(
results
) # Decrement count by the number of results fetched in this page.
start_index += 10 # Increment start index for the next page
else:
break # No more results from Google PSE, break the loop
if filter_list:
all_results = get_filtered_results(all_results, filter_list)
return [
SearchResult(
link=result["link"],
title=result.get("title"),
snippet=result.get("snippet"),
)
for result in all_results
]
```
So many people have been asking for this fix - including me, having to run the setup behind http proxy.
But such a simple fix - and yet not resolved.
Btw: all search engines fail behind http proxy since the engines just give the results urls and there is a common py. That fails the sites scraping.
@ips972 commented on GitHub (Feb 16, 2025):
So many people have been asking for this fix - including me, having to run the setup behind http proxy.
But such a simple fix - and yet not resolved.
Btw: all search engines fail behind http proxy since the engines just give the results urls and there is a common py. That fails the sites scraping.
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Originally created by @kulukami on GitHub (Feb 16, 2025).
Feature Request
Now websearching is using
requestsraw requests. And it is quite useful if adding httpProxy option.https://github.com/open-webui/open-webui/blob/main/backend/open_webui/retrieval/web/google_pse.py#L46
@ips972 commented on GitHub (Feb 16, 2025):
So many people have been asking for this fix - including me, having to run the setup behind http proxy.
But such a simple fix - and yet not resolved.
Btw: all search engines fail behind http proxy since the engines just give the results urls and there is a common py. That fails the sites scraping.