Quickstart
Your first completion in two minutes.
1. Get an API key
Create a key in the dashboard →
2. Make a request
Set SARA_API_KEY in your environment, then call the chat endpoint:
curl https://api.sara-ai.kz/v1/chat/completions \
-H "Authorization: Bearer $SARA_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gemma-4-31b",
"messages": [{"role": "user", "content": "Hello!"}],
"stream": true
}'3. Enable thinking (optional)
Gemma and GLM can think step by step before producing an answer. This behavior is controlled by the reasoning request field. When using the Python SDK, pass it in extra_body.
gemma-4-31b: thinking is disabled by default. To enable it, send"reasoning": {"enabled": true}.glm-5.2-744b: thinking is enabled by default. To disable it, send"reasoning": {"enabled": false}. This reduces response time and cost.glm-5.2-744balso supports effort levels:low,mediumandhigh. For example,"reasoning": {"effort": "low"}requests shorter thinking. The effort level is advisory and does not limit token usage; to set a hard limit, usemax_tokens.
curl https://api.sara-ai.kz/v1/chat/completions \
-H "Authorization: Bearer $SARA_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gemma-4-31b",
"messages": [{"role": "user", "content": "Is 9.11 bigger than 9.9?"}],
"reasoning": {"enabled": true}
}'
# thinking: choices[0].message.reasoning_content
# answer: choices[0].message.contentThe thinking is returned in reasoning_content, separately from the answer in content. In streaming mode, it is delivered in delta.reasoning_content.
Thinking tokens are billed as output tokens. If you set max_tokens, allow sufficient headroom: with a low limit, the model may spend the entire budget on thinking and return an empty answer.