model identifier to send to POST /v1/responses.
Sections on this page: At a glance (pricing table), Detailed model cards, and Model library by task.
At a glance
Detailed model cards
deepseek-v4-flash
1,048,576 context window$0.16 / 1M input$0.38 / 1M output$0.006 / 1M cached input
DeepSeek’s DeepSeek-V4-Flash is an open-weight Mixture-of-Experts model built for efficient reasoning, coding, and agentic workflows, with 284B total parameters activating only 13B per token. Its…gpt-oss-120b
131,072 context window$0.15 / 1M input$0.60 / 1M output
OpenAI’s gpt-oss-120b is an open-weight Mixture-of-Experts model with 117B total parameters (5.1B active per token), served on ZeroGPU for general text generation. It reasons through a problem…qwen3-30b-a3b-fp8
32,768 context window$0.05 / 1M input$0.30 / 1M output
Alibaba’s Qwen3-30B-A3B is an open-weight Mixture-of-Experts model with 30.5B total parameters (3.3B active per token), served on ZeroGPU as an FP8 build for efficient inference. It thinks through a problem…glm-5.2
262,144 context window$1.10 / 1M input$3.50 / 1M output$0.40 / 1M cached input
Z.ai’s GLM-5.2 is an open-weight Mixture-of-Experts flagship built for long-horizon tasks, with 753B total parameters activating 8 of 256 experts per token. It sustains a solid 262,144-token (256K)…llama-3.1-8b-instruct-fast
131,072 max tokens$0.02 / 1M input$0.05 / 1M output
Meta’s Llama 3.1 Instruct, tuned for fast, low-cost summarization at scale on the ZeroGPU edge network. Its 128K-token context window takes in entire documents, long transcripts, and full email or…zlm-v2-iab-classify-edge-enriched
800 max tokens$0.02 / 1M input$0.05 / 1M output
The enriched variant of ZeroGPU’s IAB classifier turns a single inference call into a full content-intelligence profile not just a label, but everything a contextual pipeline needs to act on. Each…zlm-v1-iab-classify-edge
400 max tokens$0.02 / 1M input$0.05 / 1M output
ZeroGPU’s IAB classifier maps any text to the industry-standard IAB Content Taxonomy in a single, fast inference call. Each call returns categories across both the 1.0 and 2.2 taxonomies plus matched…zlm-v1-iab-domain-classifier
100 max tokens$0.02 / 1M input$0.05 / 1M output
ZeroGPU’s Domain IAB Classifier maps a raw domain name straight to the IAB Content Taxonomy, returning content categories, topics, keywords, and user-intent signals. It needs only the domain as input, cutting payload size by up to 10x versus page-level…zlm-v1-moderation-edge
800 max tokens$0.02 / 1M input$0.05 / 1M output
ZeroGPU’s moderation model returns the complete OpenAI 13-category taxonomy — a flagged verdict, category booleans, and calibrated scores — as a drop-in for omni-moderation-latest. In head-to-head benchmarks it wins the binary safe/unsafe decision (0.899 vs 0.853 F1) and 9 of 13 harm categories, and returns verdicts 1.2–1.8× faster on production-range…gliner-multi-pii-v1
800 max tokens$0.02 / 1M input$0.05 / 1M output
GLiNER Multi PII is a multilingual PII detection and redaction model that supports on-prem deployments as well. It identifies 40+ personally identifiable entity types — identity, contact, government…gliner2-base-v1
800 max tokens$0.02 / 1M input$0.05 / 1M output
gliner2-base-v1 is a versatile extraction-and-classification model for the structured tasks that fill most production pipelines. Point it at any text and, with a single API call, pull named entities…deberta-v3-small
400 max tokens$0.02 / 1M input$0.05 / 1M output
Microsoft’s DeBERTa-v3-small is a fast, lightweight zero-shot text classifier for high-volume routing, filtering, and tagging. Hand it any text alongside your own candidate labels and it returns a…LFM2.5-1.2B-Thinking
32,768 max tokens$0.02 / 1M input$0.05 / 1M output
Liquid AI’s LFM2.5-1.2B-Thinking is a compact reasoning model that works through a problem step by step before it answers. Built by Liquid AI, it generates an explicit chain-of-thought trace so for…LFM2.5-1.2B-Instruct
32,768 max tokens$0.02 / 1M input$0.05 / 1M output
Liquid AI’s LFM2.5-1.2B-Instruct is a hybrid architecture model purpose-built for on-device deployment, trained on 28 trillion tokens with multi-stage reinforcement learning. It delivers…Model library by task
Open Weight
5 models:
deepseek-v4-flash, glm-5.2, qwen3-30b-a3b-fp8, gpt-oss-120b, llama-3.1-8b-instruct-fastText Generation
7 models:
deepseek-v4-flash, gpt-oss-120b, qwen3-30b-a3b-fp8, glm-5.2, llama-3.1-8b-instruct-fast, LFM2.5-1.2B-Thinking, LFM2.5-1.2B-InstructText Classification
4 models:
zlm-v2-iab-classify-edge-enriched, zlm-v1-iab-classify-edge, zlm-v1-iab-domain-classifier, deberta-v3-smallModeration
1 model:
zlm-v1-moderation-edgeData Extraction
2 models:
gliner-multi-pii-v1, gliner2-base-v1Ad Tech
3 models:
zlm-v1-iab-classify-edge, zlm-v2-iab-classify-edge-enriched, zlm-v1-iab-domain-classifier
