Clarified how max_output_tokens interacts with thought tokens and thinking_level on the Thinking and Generate content — Thinking pages: it is a hard infrastructure cutoff (billing still applies for any thinking tokens generated); lower t…
Changelog
Thinking
Clarified how max_output_tokens interacts with thought tokens and thinking_level on the Thinking and Generate content — Thinking pages: it is a hard infrastructure cutoff (billing still applies for any thinking tokens generated); lower thinking_level instead of a small max_output_tokens to reduce cost or latency without truncating responses.