Skip to content

types/client: relax think type to support model-defined thinking levels - #744

Merged
ParthSareen merged 2 commits into
ollama:mainfrom
arrase:fix/relax-think-type
Sep 29, 2026
Merged

ParthSareen merged 2 commits into
ollama:mainfrom
arrase:fix/relax-think-type

Conversation

@arrase

@arrase arrase commented Sep 20, 2026

Copy link
Copy Markdown
Contributor

Closes #743

What & Why

Models like Qwen 3.8 advertise custom thinking levels (such as 'xhigh') through Ollama's /api/show thinking capabilities metadata (introduced in ollama/ollama#18473), and the Ollama Go backend natively accepts any boolean or string for ThinkValue (types/model/thinking.go).

Previously, think was constrained to Optional[Union[bool, Literal['low', 'medium', 'high']]] across request models and client methods. Because ChatRequest and GenerateRequest are Pydantic models validated before requests are dispatched, passing valid model-defined levels like 'xhigh' (or 'max') resulted in a Pydantic ValidationError:

pydantic_core._pydantic_core.ValidationError: 2 validation errors for ChatRequest
think.bool
  Input should be a valid boolean, unable to interpret input [type=bool_parsing, input_value='xhigh', input_type=str]
think.literal['low','medium','high']
  Input should be 'low', 'medium' or 'high' [type=literal_error, input_value='xhigh', input_type=str]

This PR relaxes the think annotation to Optional[Union[bool, str]] = None across all 14 annotation sites in _types.py and _client.py, matching the Go backend and allowing any valid model-defined string thinking level while preserving full backward compatibility.

Changes

  • ollama/_types.py: Updated GenerateRequest.think and ChatRequest.think from Literal['low', 'medium', 'high'] to str.
  • ollama/_client.py: Updated all 12 generate and chat overloads and implementations in Client and AsyncClient to match.
  • tests/test_client.py:
    • Updated test_generate_think_annotation_matches_chat to verify annotation parity across Client.chat, Client.generate, AsyncClient.chat, and AsyncClient.generate.
    • Added test_client_chat_with_think_level verifying client.chat serialization and response handling with string thinking levels (think='xhigh').
  • tests/test_type_serialization.py:
    • Added test_think_model_defined_levels_serialization verifying serialization of model-defined levels ('xhigh', 'max', 'low', 'medium', 'high').
    • Added test_think_boolean_serialization verifying boolean thinking levels.

Testing

All 104 tests pass:

tests/test_client.py ...................................................
tests/test_type_serialization.py ....................
tests/test_utils.py ..........
104 passed in 1.33s

Linter and formatter check clean:

ruff check .
All checks passed!

arrase and others added 2 commits September 21, 2026 01:46
Models like Qwen 3.8 advertise custom thinking levels (such as 'xhigh')
through /api/show thinking metadata, and the Ollama Go backend
accepts any bool or string for ThinkValue.

Previously, think was restricted to Optional[Union[bool, Literal['low', 'medium', 'high']]],
causing Pydantic ValidationError in ChatRequest and GenerateRequest when passing
supported levels like 'xhigh' or 'max'. Relaxing the annotation to
Optional[Union[bool, str]] allows all valid model-defined thinking levels.
@ParthSareen
ParthSareen self-requested a review September 29, 2026 01:13
@ParthSareen
ParthSareen merged commit 8785556 into ollama:main Sep 29, 2026
2 checks passed
@ParthSareen

Copy link
Copy Markdown
Member

Thank you for this!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

ChatRequest and GenerateRequest reject model-defined thinking levels (e.g. 'xhigh') with Pydantic ValidationError

2 participants