All models

GLM-5.3 Flash runs at 47% off on Moyi API. .

ZhipuGLM: GLM-5.3 Flash

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead.

Modalities

Price (in / out)

47% off

$0.06 / $0.2104$0.114 / $0.4per 1M

Context

1M

Released

Aug 26, 2026

Providers

Moyi API routes by health and cost within the discount range and provider list you configured. When one returns an error we fall through to the next best — you don't pick a provider in the request. Configure discount range

Need a specific set of providers?

ProviderContextInputOutputCacheLatencyThroughputNo training
Z.AI1M$0.114$0.06$0.4$0.2104$0.03$0.0158
VolcengineVolcengine1M$0.114$0.06$0.4$0.2104$0.03$0.0158
OpenRouterOpenRouter1M$0.114$0.06$0.4$0.2104$0.03$0.0158
Third-party providers1M$0.114$0.06$0.4$0.2104$0.03$0.0158

Live discount

Each provider's average effective discount, refreshed hourly.

No discount data for this model in the last 24 hours.

Performance

Latency and throughput are measured on live gateway traffic.

Latency

Time from sending the request to the first output token, broken down by provider.

    LATENCY · MOYI API GATEWAY

    Throughput

    Tokens per second after the first output token, broken down by provider.

      THROUGHPUT · MOYI API GATEWAY

      Uptime

      Moyi API continuously probes every provider and falls through to the next best one on error.

      Uptime (24h)

      Uptime (3d)

      Uptime, last 3 days

      -72h-48h-24hNow

      Uptime, last 24 hours

      One line per upstream provider, taking the best success rate among that provider's channels. The Moyi API line counts only the final attempt of each request — the end-to-end result after all fallbacks.

        AVAILABILITY · LAST 24H · MOYI API GATEWAY

        Data security

        Among this model's providers, no provider do not train on your data.

        For the full data-security policy, learn more.

        Need to pick your own providers? .

        Quickstart

        from openai import OpenAI
        
        client = OpenAI(
            base_url="https://api.moyiapi.com/v1",  # the only line that changes
            api_key="<MOYIAPI_API_KEY>",
        )
        response = client.chat.completions.create(
            model="glm-5.3-flash",
            messages=[{"role": "user", "content": "Hello"}],
        )
        print(response.choices[0].message.content)
        Read docs

        Endpoint: https://api.moyiapi.com/v1

        GLM-5.3 Flash pricing, performance & uptime