首页/GLM 5.3 Flash
ChatGLM

GLM 5.3 Flash

zai-org/glm-5.3-flash
GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead.
Price
Input$0.15 per million tokens
Cached reads$0.03 per million tokens
Output$0.5 per million tokens

Use the following code example to integrate our API:

1from openai import OpenAI
2
3client = OpenAI(
4    api_key="<Your API Key>",
5    base_url="https://api.highwayapi.ai/openai"
6)
7
8response = client.chat.completions.create(
9    model="zai-org/glm-5.3-flash",
10    messages=[
11        {"role": "system", "content": "You are a helpful assistant."},
12        {"role": "user", "content": "Hello, how are you?"}
13    ],
14    max_tokens=131072,
15    temperature=0.7
16)
17
18print(response.choices[0].message.content)

Information

Provider
Quantification
fp8

Supported Features

Context length
1M
Maximum output
131.1K
Function call
Support
Structured output
Support
Reasoning
Support
serverless
Support
Input Capabilities
text, image
Output Capabilities
text
Contact Us