ollama

mirror of https://github.com/ollama/ollama.git synced 2026-03-12 01:45:29 -05:00

Author	SHA1	Message	Date
Parth Sareen	35c3c9e3c2	anthropic: allow non-thinking models when using Anthropic API (#13692 )	2026-01-12 15:13:26 -08:00
Jeffrey Morgan	9667c2282f	x/imagegen: add naive TeaCache and FP8 quantization support (#13683 ) TeaCache: - Timestep embedding similarity caching for diffusion models - Polynomial rescaling with configurable thresholds - Reduces transformer forward passes by ~30-50% FP8 quantization: - Support for FP8 quantized models (8-bit weights with scales) - QuantizedMatmul on Metal, Dequantize on CUDA - Client-side quantization via ollama create --quantize fp8 Other bug fixes: - Fix `/api/show` API for image generation models - Server properly returns model info (architecture, parameters, quantization) - Memory allocation optimizations - CLI improvements for image generation	2026-01-12 13:45:22 -08:00
Jeffrey Morgan	a937a68317	server: fix slow 'ollama rm' of models with many layers (#13680 ) RemoveLayers was calling Manifests() for each layer to check if it was shared with other models. For models with many blobs (e.g., tensor models), this caused O(N*M) manifest reads. Now loads manifests once and builds a set of in-use digests.	2026-01-12 13:17:48 -08:00
Jeffrey Morgan	2584940016	Add z-image image generation prototype (#13659 )	2026-01-09 21:09:46 -08:00
Parth Sareen	2b2cda7a2b	api: implement anthropic api (#13600 ) * api: add Anthropic Messages API compatibility layer Add middleware to support the Anthropic Messages API format at /v1/messages. This enables tools like Claude Code to work with Ollama local and cloud models through the Anthropic API interface.	2026-01-09 11:53:36 -08:00
Daniel Hiltgen	33ee7168ba	Add experimental MLX backend and engine with imagegen support (#13648 ) * WIP - MLX backend with gemma3 * MLX: add cmake and go tag build toggles To build the new MLX backend code: cmake --preset MLX cmake --build --preset MLX --parallel cmake --install build --component MLX go build -tags mlx . Note: the main.go entrypoint for the MLX engine will change in a follow up commit. * add experimental image generation runtime * add experimental image generation runtime * MLX: wire up cuda build for linux * MLX: get dependencies correct and dedup This is still too large for a unified github artifact, but is now "correct" for the mlx_cuda_v13 directory. * fix relative link bug in dedup * Add darwin build and readme * add go build tag for mlx dependent code and wire up build_darwin.sh * lint cleanup * macos: build mlx for x86 This will be CPU only. * cuda build instructions and fix drift from mlx bump * stale comment * Delete agent helper doc * Clean up readme.md * Revise README for tokenizer clarity and details Updated README to clarify tokenizer functionality and removed correctness section. --------- Co-authored-by: jmorganca <jmorganca@gmail.com>	2026-01-08 16:18:59 -08:00
Devon Rifkin	e51dead636	preserve tool definition and call JSON ordering (#13525 ) * preserve tool definition and call JSON ordering This is another iteration of <https://github.com/ollama/ollama/pull/12518>, but this time we've simplified things by relaxing the competing requirements of being compatible AND order-preserving with templates (vs. renderers). We maintain backwards compatibility at the cost of not guaranteeing order for templates. We plan on moving more and more models to renderers, which have been updated to use these new data types, and additionally we could add an opt-in way of templates getting an order-preserved list (e.g., via sibling template vars) * orderedmap_test: remove testify	2026-01-05 18:03:36 -08:00
lif	37f6f3af24	server: return error when embedding contains NaN or Inf values (#13599 ) The normalize function now checks for NaN and Inf values in the embedding vector before processing. This prevents JSON encoding failures when models produce invalid floating-point values. Fixes #13572 Signed-off-by: majiayu000 <1835304752@qq.com>	2026-01-03 02:20:12 -05:00
Jeffrey Morgan	8852220f59	add REQUIRES command to Modelfile (#13361 )	2025-12-18 13:21:29 -08:00
Grace	522c11a763	Revert "Omit args and params in tool function def and calls (#13516 )" (#13518 ) This reverts commit `0fadeffaee`.	2025-12-17 19:06:56 -08:00
Grace	0fadeffaee	Omit args and params in tool function def and calls (#13516 )	2025-12-17 18:42:21 -08:00
Bruce MacDonald	45c4739374	types: ConfigV2 and RootFS (#13504 ) Refactored the ConfigV2 and RootFS types from server/images.go to a new types/model/config.go file under the model package. Updated all references to use model.ConfigV2 and model.RootFS. This allows for use in other projects without worrying about compiling the c code in the llama package.	2025-12-16 15:18:17 -08:00
Devon Rifkin	1eb5e75972	openai: add v1/responses support (#13351 ) Only supporting the stateless part of the API. Doc updates to come once this is shipped. Closes: #9659	2025-12-11 15:37:10 -08:00
EasonLin	1c4e85b4df	routes: add logprobs in tool calls (#13238 )	2025-12-10 17:28:41 -08:00
nicole pardal	e082d60a24	truncation: fixed runner truncation logic + removed server truncation (#12839 ) This PR consolidates all embedding prompt-length checking, truncation, and prompt token counting into the runner to ensure a single source of truth.	2025-12-08 11:20:28 -08:00
Sos Pogosyan	31b8c6a214	fix(api): correct Content-Type header for /api/chat and /api/generate when using cloud models (#13279 ) --------- Co-authored-by: Pogosyan Sos <sos_pogosyan@MacBook-Pro-Sos.local> Co-authored-by: Patrick Devine <patrick@infrahq.com>	2025-12-04 21:33:07 -08:00
Grace	d70e935526	Parser for Cogito v2 (#13145 )	2025-11-19 17:21:07 -08:00
Michael Yang	718961de68	migrate to golangci-lint v2 (#13109 ) * migrate to golangci-lint v2 * copyloopvar	2025-11-18 11:00:26 -08:00
omahs	dd0ed0ef17	docs: fix typos in repository documentation (#10683 )	2025-11-15 20:22:29 -08:00
pierwill	4cea757e70	server: clean up manifest documentation (#12995 ) Co-authored-by: pierwill <pierwill@users.noreply.github.com>	2025-11-15 19:13:15 -08:00
Parth Sareen	c114987523	logprob: add bytes to logprobs (#13068 )	2025-11-13 13:49:25 -08:00
Jesse Gross	f560bd077f	llm: Use Ollama engine memory layouts for both old and new engines Currently for both the old and new engines, there is code to calculate how much memory is required for a model and lay out the layers onto GPUs. This reuses the new engine's lay out code for the old engine as well, bringing them closer together. The old engine continues to use its current method of estimating required memory. This reduces maintainence effort and improves consistency, as new features only need to be implemented in one place. The newer code is also more accurate, especially with multiple GPUs.	2025-11-11 13:11:08 -08:00
Baptiste Jamin	59241c5bee	server: add logprobs and top_logprobs support to Ollama's API (#12899 ) Adds logprobs support to Ollama's API including support for Ollama's OpenAI-compatible API. By specifying the new 'logprobs' boolean parameter in the API, Ollama will return the log probabilities for each token generated. 'top_logprobs', an integer value can also be specified up to the value 20. When specified, the API will also provide the number of most likely tokens to return at each token position Co-authored-by: Baptiste Jamin <baptiste@crisp.chat>	2025-11-11 08:49:50 -08:00
breatn	780762f9d2	server: fix duplicate 'is' typo in comment (#12985 )	2025-11-06 14:44:44 -08:00
Jeffrey Morgan	30fcc71983	api: add omitempty to required tool function parameter type (#12989 )	2025-11-06 14:08:55 -08:00
Daniel Hiltgen	408c2f99d0	log: trace logging for scheduler (#12961 )	2025-11-05 08:12:15 -08:00
Grace	809b9c68fa	Add Tool Call ID (#12956 ) * routes/types: add tool call id --------- Co-authored-by: ParthSareen <parth.sareen@ollama.com>	2025-11-04 16:43:33 -08:00
Daniel Hiltgen	d3b4b9970a	app: add code for macOS and Windows apps under 'app' (#12933 ) * app: add code for macOS and Windows apps under 'app' * app: add readme * app: windows and linux only for now * ci: fix ui CI validation --------- Co-authored-by: jmorganca <jmorganca@gmail.com>	2025-11-04 11:40:17 -08:00
Michael Yang	7d25b9e194	feat(model): add qwen3vl (#12665 )	2025-10-28 17:39:47 -07:00
Patrick Devine	29f63f37c8	Revert "server: Consolidate embedding truncation in runner (#12730 )" (#12810 ) This reverts commit `5d347f6d6f`.	2025-10-28 14:49:14 -07:00
Devon Rifkin	9862317174	Merge pull request #12793 from ollama/drifkin/12792_renderer-parser-from create: inherit FROM model's renderer/parser	2025-10-28 00:15:46 -07:00
Devon Rifkin	1bdd816910	create: inherit FROM model's renderer/parser On main, the `RENDERER` and `PARSER` fields from the `Modelfile` don't get propagated to a new model created with a `req.From` parameter. This is easily triggered via `ollama run qwen3-coder`, then running some save command like `/save qwen3-coder-custom`. Added a regression test for this, and then open the config for the "from" model in order to use its renderer/parser as a default for the new model. This will fix the CLI and also API-based creates. Fixes: https://github.com/ollama/ollama/issues/12792	2025-10-27 15:14:19 -07:00
nicole pardal	5d347f6d6f	server: Consolidate embedding truncation in runner (#12730 ) Currently, checking the length of prompts for embeddings to ensure they fit in the context window (and possible truncation) occurs in two places - the Ollama server and runner. This can lead to inconsistencies in both the checks and reported number of tokens processed. Since we have to do this processing in the runner, this consolidates all of the logic there.	2025-10-27 11:59:12 -07:00
Patrick Devine	b97eb2b858	cloud: set the proxy content-type to the same as local models (#12759 )	2025-10-25 10:57:10 -07:00
Daniel Hiltgen	3258a89b6e	DRY out the runner lifecycle code (#12540 ) * DRY out the runner lifecycle code Now that discovery uses the runners as well, this unifies the runner spawning code into a single place. This also unifies GPU discovery types with the newer ml.DeviceInfo * win: make incremental builds better Place build artifacts in discrete directories so incremental builds don't have to start fresh * Adjust sort order to consider iGPUs * handle cpu inference oom scenarios * review comments	2025-10-23 11:20:02 -07:00
Patrick Devine	d515aed6c3	cloud: don't error sending empty messages (#12724 )	2025-10-21 18:12:14 -07:00
Michael Yang	d2b63c19b3	fs(ggml): fill in arch prefix if necessary (#12646 )	2025-10-20 16:42:18 -07:00
Daniel Hiltgen	68e04c7ff8	test: harden scheduler tests (#12662 ) * test: harden scheduler tests This removes reschedDelay which was stale code, and adds a new configurable timeout for the waitForVRAMRecovery so tests can now set the timeout to be very short to avoid the scheduler getting stuck and hitting a test timeout. * test: tune tests for partial loads Give stress tests more time when the model is split between CPU/GPU	2025-10-17 08:56:44 -07:00
Jeffrey Morgan	65fb3ff49d	renderers: add global flag for setting [img] tags (#12669 ) Adds a temporary global flag to renderers that causes renderers to always render images as [img]. In a follow up change, we will consider making this the default, and this flag could eventually be removed	2025-10-16 16:37:32 -07:00
Devon Rifkin	ddaca643d0	add registries for parsers/renderers	2025-10-14 01:13:54 -07:00
Grace	05982a95cb	Qwen3VL Cloud Parser and Renderer (#12526 ) * working (other than tool call is the incorrect order) for tool calls and tools * Tests work, other than image tags (tests do not go through server) and tools (not in the correct order, but contents are the same) * testing for qwen3vl parser - toolparser is working * made changes to JSON tool parser, wraps the TollCallFunction with a TollCall object * Working parser for thinking models - assumes state of thinking, emits unambiguous content in thinking, does not call tool call in thinking * changed the parser to start with collecting content * thinking prefill * add hasThinkingSupport parameter to parser * qwen3-vl -> qwen3-vl-instruct for renderer/parser * Add hasThinkingSupport=false to QwenVLParser --------- Co-authored-by: Devon Rifkin <drifkin@drifkin.net>	2025-10-13 16:52:33 -07:00
Jeffrey Morgan	6544e14735	Reapply "add truncate and shift parameters" (#12582 )	2025-10-11 16:06:14 -07:00
Devon Rifkin	6db8da9958	routes: fix built-in renderers for `api/generate` Made it so when api/generate builds up a message array and generates the prompt it now goes through the same function as `api/chat` for consistency. This is where we hook the optional built-in renderers to bypass templates, which was missing for `api/generate` before this change. Closes: #12578	2025-10-11 13:57:43 -07:00
Daniel Hiltgen	aab2190420	implement nvml for linux (#12517 ) * implement nvml for linux * Improve scheduler logging when VRAM doesn't recover	2025-10-10 15:15:56 -07:00
Patrick Devine	d681cd7c29	thinking: allow `"think": false` for non-thinking models (#12555 )	2025-10-09 18:46:00 -07:00
Daniel Hiltgen	15e3611d3d	logs: quiet down context canceled on completion and scheduler noise (#12553 ) * logs: quiet down context canceled on completion If the client closes the connection before Completion finishes, we were logging at error level implying the runner crashed which was misleading. time=2025-10-08T22:59:20.566-07:00 level=ERROR source=server.go:1490 msg="post predict" error="Post \"http://127.0.0.1:57736/completion\": context canceled" * quiet down scheduler log error on expected case Since we don't hold the lock while performing memory load calculations, other runners can unload in parallel, so finding no runner to unload is a valid scenario which we shouldn't log at error level.	2025-10-09 10:37:47 -07:00
Parth Sareen	77060d462c	routes: structured outputs for gpt-oss (#12460 )	2025-10-08 19:13:38 -07:00
Jeffrey Morgan	7d965258ce	Revert "add truncate and shift parameters (#12519 )" (#12545 ) This reverts commit `6a62b894c7`.	2025-10-08 17:57:57 -07:00
Jeffrey Morgan	6a62b894c7	add truncate and shift parameters (#12519 )	2025-10-08 17:05:05 -07:00
Patrick Devine	90d429f5a8	thinking: turn on thinking mode for all reasoning models (#12533 )	2025-10-08 16:50:13 -07:00

1 2 3 4 5 ...

953 Commits