All Z.ai GLM API Plans & Pricing

Plan Monthly Annual Best For
View all features by plan (compare side-by-side)

GLM-5.2 ($1.40 in / $4.40 out per 1M tokens)

  • $1.40 per 1M input tokens
  • $0.26 cached input
  • $4.40 per 1M output tokens

GLM-5 ($1.00 in / $3.20 out per 1M tokens)

  • $1.00 per 1M input tokens
  • $0.20 cached input
  • $3.20 per 1M output tokens

GLM-4.7 ($0.60 in / $2.20 out per 1M tokens)

  • $0.60 per 1M input tokens
  • $0.11 cached input
  • $2.20 per 1M output tokens

GLM-4.7-Flash / GLM-4.5-Flash (free)

  • Free model tier
  • GLM-4.6V-Flash vision also free
Pricing Alerts

Track Z.ai GLM API pricing

Get an email when Z.ai GLM API's pricing changes — plus the weekly SaaS Price Watch: verified price changes and deals across 3,000+ products. One-click unsubscribe.

Compare Z.ai GLM API with alternativesAdjust seats, lock a tier, add up to 2 more products side-by-side. Shareable URL.
Quick Answer
Last verified:
High confidence

Z.ai GLM API costs Free to $4.40 per million tokens as of August 2026, with 4 plans available including a free tier. Plans: GLM-5.2 ($1.40 in / $4.40 out per 1M tokens) (free), GLM-5 ($1.00 in / $3.20 out per 1M tokens) (free), GLM-4.7 ($0.60 in / $2.20 out per 1M tokens) (free), and GLM-4.7-Flash / GLM-4.5-Flash (free) (free). Pricing depends on your chosen tier, contract length, and negotiated discounts.

Use the interactive pricing calculator to estimate your exact cost based on team size and requirements.

  • Free tier: Yes

Z.ai GLM API offers 4 pricing tiers: GLM-5.2 ($1.40 in / $4.40 out per 1M tokens), GLM-5 ($1.00 in / $3.20 out per 1M tokens), GLM-4.7 ($0.60 in / $2.20 out per 1M tokens), GLM-4.7-Flash / GLM-4.5-Flash (free). The GLM-5 ($1.00 in / $3.20 out per 1M tokens) plan is strong general-purpose workloads.

Compared to other llm api providers software, Z.ai GLM API is positioned at the budget-friendly price point.

  • 10 documented hidden costs beyond list price

How much does Z.ai GLM API cost?

Z.ai GLM API offers a free plan with 3 paid tiers from $Infinity to $4.40 per million tokens. Plans include GLM-5.2 ($1.40 in / $4.40 out per 1M tokens) (free), GLM-5 ($1.00 in / $3.20 out per 1M tokens) (free), GLM-4.7 ($0.60 in / $2.20 out per 1M tokens) (free), GLM-4.7-Flash / GLM-4.5-Flash (free) (free).

Z.ai GLM API Pricing Overview

Z.ai GLM API has 4 pricing plans, including a free tier. Paid plans range from $0 to $4.40/per million tokens. The GLM-5.2 ($1.40 in / $4.40 out per 1M tokens) plan is free and is best for frontier-quality glm workloads. The GLM-5 ($1.00 in / $3.20 out per 1M tokens) plan is free and is best for strong general-purpose workloads. The GLM-4.7 ($0.60 in / $2.20 out per 1M tokens) plan is free and is best for cost-efficient production workloads. The GLM-4.7-Flash / GLM-4.5-Flash (free) plan is free and is best for prototyping and light workloads at zero cost.

There are at least 10 documented hidden costs beyond Z.ai GLM API's list price, including implementation, training, and add-on fees.

This pricing was last verified in August 17, 2026 from 1 independent source.

How Z.ai GLM API Pricing Compares

Compare Z.ai GLM API pricing against top alternatives in LLM API Providers.

Compare Z.ai GLM API vs Alternatives

Before committing to Z.ai GLM API, compare pricing with these 3 alternatives in the same category.

All Z.ai GLM API alternatives & migration guides

What Companies Actually Pay for Z.ai GLM API

Review scores
Trustpilot 2.2out of 5 (31)
Third-party review aggregates, as of Aug 2026
Top pricing complaints
Strict Usage Limits and High CostUnreliability and Slow PerformancePoor Inference PerformanceLack of Customer Support and Refunds

How Z.ai GLM API Pricing Compares

Software Starting Price Top Price
Z.ai GLM API Free $4.4 per million tokens
Amazon Bedrock $0.07 per million tokens $75 per million tokens
Anyscale Free $4.9591 per million tokens
Baidu ERNIE API $0.1 per million tokens $10 per million tokens
Cerebras Inference API $0.1 per million tokens $6 per million tokens
Cohere API Free $15 per million tokens

10 Z.ai GLM API Hidden Costs Beyond the List Price

Beyond the listed price, Z.ai GLM API has at least 10 documented hidden costs that can significantly increase total cost of ownership.

Watch for 10 hidden costs
  • Cached Input Storage $0.26 per 1 million tokens
    medium 1 source
    industry "Cached Input Storage: Z.ai currently lists cached input storage as free for a limited time, implying it could become a paid feature in the future"
  • Premium Model Tiers 64% higher for input and 81% higher for output
    high 1 source
    industry "Its high-throughput tier, GLM-5.2-Fast, is priced higher at $2.29 per 1 million input tokens and $8.00 per 1 million output tokens"
  • Conversation History Reprocessing 10x cost multiplier
    critical 1 source
    industry "A 50-message thread can send the equivalent of a short document, leading to a "10x cost multiplier" if not managed through prompt caching"
  • Suboptimal Model Selection
    medium 1 source
    industry "For instance, GLM-5.2, a flagship model, costs $1.40 per 1 million input tokens and $4.40 per 1 million output tokens"
  • Data Training Opt-ins
    high 1 source
    industry "Cheaper models like GLM-4.7-FlashX are available at $0.07 input and $0.40 output per 1 million tokens, with some "Flash" models being free"
  • System Prompt Overhead $400 monthly
    high 1 source
    industry "A 2,000-token system prompt with 100,000 requests per month could cost around $400 monthly at GPT-4.1 rates before any user tokens are counted."
  • Uncapped Output Generation
    critical 1 source
    industry "Uncapped Output Generation: Models can produce variable-length responses, potentially exceeding projected output tokens by 10-25 times if max_tokens limits are not explicitly set."
  • Web Search Tool Usage $0.01 per use
    low 1 source
    industry "ai's built-in Web Search tool costs an additional $0.01 per use, on top of the token charges for the request."
  • Latency and Data Residency
    medium 1 source
    industry "ai's infrastructure is primarily located in China, which could affect latency for international users and raise data residency concerns, potentially leading to additional compliance or infrastructure costs for some buyers."
  • Model Verbosity
    medium 1 source
    industry "Model Verbosity: Even with identical per-token rates, a more verbose model can lead to higher overall costs for a completed workload due to generating more output tokens."
Tip

Ask your Z.ai GLM API sales rep about these costs upfront. Getting them in writing before signing can save you from surprise charges later.

Full hidden costs breakdown →

Intelligence sourced from 2 independent sources
industry Trustpilot Consumer reviews
Key claims include inline source attribution. Data verified against multiple independent sources. 10 source citations total.

Is this pricing incorrect? — we'll verify and update it.