$ yogi.compute

Creator rewards, spent as compute. Hold 100,000 $YOGI and the API turns on. Sell, and it turns off. There is nothing to buy and no balance to top up.

Inference

Token

01

Every trade pays a fee

$YOGI generates creator fees on every trade. Those fees pay for inference.

02

Holding is the subscription

Hold 100,000 $YOGI and the API turns on. Sell, and it turns off.

03

Then use it, unmetered

Not a quota and not a share. Every holder gets the same access to all models.

Every model, unmetered

Prices are what each one costs the treasury, not what it costs you.

/
104 models
AnthropicAnthropic

Claude Opus 5

1M

Anthropic's flagship model for demanding reasoning, coding, and long-horizon agentic work.

$5/M in · $25/M outRun now →
OpenAIOpenAI

GPT-6 Astra

1.1M

OpenAI's flagship model for demanding end-to-end work: analysis, software engineering, deep research.

$10/M in · $50/M outRun now →
AnthropicAnthropic

Claude Fable 5.1

1M

Mythos-class model with the biggest gains in agentic coding and long-running workflows.

$10/M in · $50/M outRun now →
SpaceXAISpaceXAI

Grok 4.6

500K

Frontier performance on coding, knowledge work, and STEM.

$2/M in · $6/M outRun now →
Z.aiZ.ai

GLM 5.3

1.3M

Large-scale reasoning model built for complex software engineering and long-horizon agent tasks.

$1.4/M in · $4.4/M outRun now →
QwenQwen

Qwen3.8 2.4T A95B

1M

Open-weight sparse mixture-of-experts model, the open-weight variant of Qwen3.8 Max.

$2/M in · $6/M outRun now →
GoogleGoogle

Gemini 3.8 Flash

1M

Google's most intelligent Flash model, with significant gains across coding and agentic tasks.

$0.75/M in · $3.75/M outRun now →
DeepSeekDeepSeek

DeepSeek V4 Pro 0813

1M

The GA release of DeepSeek V4 Pro.

$1.05/M in · $3.15/M outRun now →
MetaMeta

Muse Spark 1.3

1M

Multimodal reasoning model from Meta for long-running agentic, multi-agent, and coding workflows.

$1.25/M in · $4.25/M outRun now →
MoonshotAIMoonshotAI

Kimi K3

1M

2.8T parameter open-weight multimodal reasoning model for complex coding and knowledge work.

$3/M in · $15/M outRun now →
OpenAIOpenAI

GPT-5.6 Sol

1.1M

Flagship model in the GPT-5.6 series, suited for complex reasoning, coding, and agentic workflows.

$2/M in · $12/M outRun now →
AnthropicAnthropic

Claude Sonnet 5

1M

Anthropic's most capable Sonnet-class model, with frontier performance across coding and agents.

$2/M in · $10/M outRun now →
GoogleGoogle

Gemini 3.1 Pro Preview

1M

Google's frontier reasoning model, delivering enhanced software engineering performance.

$2/M in · $12/M outRun now →
OpenAIOpenAI

GPT-5.4

1.1M

OpenAI's latest frontier model with a 1M+ token context window.

$2.5/M in · $15/M outRun now →
DeepSeekDeepSeek

DeepSeek V4 Flash

1M

Efficiency-optimized Mixture-of-Experts model with 284B total and 13B activated parameters.

$0.089/M in · $0.177/M outRun now →
QwenQwen

Qwen3.8 Flash

1M

Multimodal reasoning model for coding assistance, agentic workflows, and visual understanding.

$0.15/M in · $0.47/M outRun now →
MistralMistral

Mistral Medium 3.5

262K

Dense 128B instruction-following model supporting text and image inputs.

$1.5/M in · $7.5/M outRun now →
MiniMaxMiniMax

MiniMax M3

1M

Multimodal foundation model with text, image, and video inputs and a 1M-token context.

$0.3/M in · $1.2/M outRun now →
SpaceXAISpaceXAI

Grok 4.20

2M

Reasoning model with industry-leading speed and agentic tool calling capabilities.

$1.25/M in · $2.5/M outRun now →
NVIDIANVIDIA

Nemotron 3.5 Lightning

262K

Open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total.

$0.08/M in · $0.2/M outRun now →
PerplexityPerplexity

Sonar Pro

200K

Perplexity's advanced search-grounded model. Pricing includes search pricing.

$3/M in · $15/M outRun now →
InceptionInception

Mercury 2.5

260K

The fastest reasoning LLM and the latest diffusion LLM, generating tokens in parallel.

$0.04/M in · $0.15/M outRun now →
ByteDance SeedByteDance Seed

Seed-2.0-Code

262K

Model optimized for agentic coding, frontend development, and multilingual programming.

$0.5/M in · $3/M outRun now →
TencentTencent

Hy4 preview

1M

Mixture-of-experts model with 49B active parameters out of 770B total, designed for coding agents.

$0.834/M in · $2.5/M outRun now →
OpenAIOpenAI

GPT-5.4 Mini

400K

Core GPT-5.4 capabilities in a faster, more efficient model for high-throughput workloads.

$0.75/M in · $4.5/M outRun now →
AnthropicAnthropic

Claude Haiku 4.5

200K

Anthropic's fastest and most efficient model, delivering near-frontier intelligence.

$1/M in · $5/M outRun now →
GoogleGoogle

Gemini 3.5 Flash Lite

1M

High-efficiency model with upgraded agentic capabilities for focused subagent workloads.

$0.3/M in · $1.25/M outRun now →
MetaMeta

Llama 4 Maverick

1M

High-capacity multimodal model from Meta, built on a mixture-of-experts architecture.

$0.2/M in · $0.696/M outRun now →
QwenQwen

Qwen3.6 27B

262K

Dense 27-billion-parameter language model with hybrid multimodal capabilities.

$0.3/M in · $2/M outRun now →
Z.aiZ.ai

GLM 4.7 Flash

203K

30B-class SOTA model balancing performance and efficiency, optimized for agentic coding.

$0.061/M in · $0.4/M outRun now →
falfal

Veo 3 Fast

video

Google's Veo 3, fast tier, with audio. The most expensive request this API serves.

$1.60 per clipRun now →
falfal

Kling 2.5 Turbo Pro

video

Kling's fast tier. Reliable prompt adherence and clean motion.

$0.35 per clipRun now →
falfal

Hailuo 02

video

MiniMax's video model. Good physical motion and camera movement.

$0.28 per clipRun now →
falfal

LTX Video 13B

video

The cheapest way to get a clip out of this API. The right default for iterating on a prompt.

$0.06 per clipRun now →
LiquidAILiquidAI

LFM2.5-2.6B (free)

66K

Compact reasoning model suited for agent workflows, data extraction, and RAG.

Free in · Free outRun now →
GoogleGoogle

Gemma 4 31B (free)

262K

Google DeepMind's 30.7B dense multimodal model with a 256K token context window.

Free in · Free outRun now →
OpenAIOpenAI

o5

400K

OpenAI's dedicated reasoning model for hard math, science, and multi-step logic.

$3/M in · $12/M outRun now →
OpenAIOpenAI

o4 Mini

200K

Compact reasoning model with strong STEM performance at a fraction of o5's cost.

$0.8/M in · $3.2/M outRun now →
OpenAIOpenAI

GPT-5.4 Nano

400K

The cheapest GPT-5.4 tier, built for high-volume classification and extraction.

$0.2/M in · $1.2/M outRun now →
OpenAIOpenAI

GPT-5.2

400K

Previous-generation frontier model, still a strong default for general assistant work.

$1.5/M in · $8/M outRun now →
OpenAIOpenAI

GPT-5.1 Codex

400K

Coding-specialized GPT-5.1 variant tuned for agentic software engineering.

$1.75/M in · $10/M outRun now →
OpenAIOpenAI

GPT-4.1

1M

Reliable 1M-context workhorse from the GPT-4 series, cheap for long-document tasks.

$1/M in · $4/M outRun now →
AnthropicAnthropic

Claude Opus 4.8

1M

Previous Opus generation with deep reasoning and careful long-form writing.

$5/M in · $25/M outRun now →
AnthropicAnthropic

Claude Sonnet 4.6

1M

Balanced Sonnet release known for dependable coding and instruction following.

$3/M in · $15/M outRun now →
AnthropicAnthropic

Claude Haiku 4

200K

Fast, inexpensive Haiku for latency-sensitive assistants and routing.

$0.8/M in · $4/M outRun now →
GoogleGoogle

Gemini 3.5 Pro

1M

Google's Pro-tier reasoning model with strong multimodal and coding scores.

$1.5/M in · $9/M outRun now →
GoogleGoogle

Gemini 3 Flash

1M

Fast Gemini for everyday chat, summarization, and light agentic work.

$0.5/M in · $2.5/M outRun now →
GoogleGoogle

Gemma 4 12B

262K

Small open Gemma model for on-device-class workloads served from the API.

$0.05/M in · $0.2/M outRun now →
GoogleGoogle

Gemma 3 27B

128K

Older open-weight Gemma, still popular for fine-tuned derivatives.

$0.04/M in · $0.15/M outRun now →
SpaceXAISpaceXAI

Grok 4

256K

The original Grok 4 flagship with strong reasoning and real-time knowledge.

$3/M in · $15/M outRun now →
SpaceXAISpaceXAI

Grok 4.1 Fast

2M

Speed-optimized Grok with a 2M context, built for agentic tool calling at scale.

$0.2/M in · $0.5/M outRun now →
SpaceXAISpaceXAI

Grok Code Fast 2

256K

Coding-tuned Grok variant for autocomplete-style and edit-heavy workflows.

$0.3/M in · $1.5/M outRun now →
MetaMeta

Llama 4 Scout

1M

Lightweight Llama 4 with a huge context window, ideal for retrieval-heavy apps.

$0.15/M in · $0.5/M outRun now →
MetaMeta

Llama 3.3 70B

128K

The classic open-weight 70B that still powers half the fine-tunes in the wild.

$0.1/M in · $0.3/M outRun now →
MetaMeta

Muse Spark 1.2

512K

Earlier Muse Spark release with solid multimodal reasoning at a lower price.

$0.9/M in · $3/M outRun now →
DeepSeekDeepSeek

DeepSeek R2

1M

DeepSeek's next-gen reasoning model with long chains of thought and tool use.

$2/M in · $8/M outRun now →
DeepSeekDeepSeek

DeepSeek V3.2

128K

The V3 refresh that made DeepSeek famous: strong chat quality, tiny price.

$0.3/M in · $0.9/M outRun now →
DeepSeekDeepSeek

DeepSeek R1

128K

The original open reasoning model that started the RL-reasoning wave.

$0.55/M in · $2.19/M outRun now →
QwenQwen

Qwen3.8 Coder

1M

Code-specialized Qwen3.8 with repo-scale context and agentic editing skills.

$1/M in · $4/M outRun now →
QwenQwen

Qwen3.6 72B

262K

Dense 72B Qwen for teams that want open weights without MoE complexity.

$0.5/M in · $2/M outRun now →
QwenQwen

Qwen3 VL 235B

262K

Vision-language flagship of the Qwen3 line, strong on documents and UI screenshots.

$0.3/M in · $1.2/M outRun now →
QwenQwen

QwQ 32B

128K

Compact open reasoning model that punches far above its size on math.

$0.15/M in · $0.6/M outRun now →
MistralMistral

Mistral Large 3

262K

Mistral's flagship dense model for enterprise assistants and RAG.

$2/M in · $6/M outRun now →
MistralMistral

Codestral 26

262K

Mistral's code model, tuned for fill-in-the-middle and completion workloads.

$0.3/M in · $0.9/M outRun now →
MistralMistral

Magistral Medium

128K

Mistral's reasoning line with transparent, inspectable chains of thought.

$2/M in · $5/M outRun now →
MistralMistral

Devstral 2

262K

Agentic coding model built with SWE-bench-style tool use in mind.

$0.4/M in · $1.2/M outRun now →
MistralMistral

Mistral Small 4

128K

Small, cheap, and surprisingly capable Mistral for everyday tasks.

$0.1/M in · $0.3/M outRun now →
MoonshotAIMoonshotAI

Kimi K2.5

262K

The K2 refresh that made Kimi a favorite for long-context agentic coding.

$0.6/M in · $2.5/M outRun now →
MoonshotAIMoonshotAI

Kimi K3 Thinking

1M

Deliberate-reasoning variant of K3 for the hardest problems, no time pressure.

$4/M in · $18/M outRun now →
Z.aiZ.ai

GLM 5

203K

Z.ai's previous flagship, still strong on coding and agentic benchmarks.

$1/M in · $3.2/M outRun now →
Z.aiZ.ai

GLM 4.7

203K

Dependable mid-tier GLM for assistants, extraction, and routing.

$0.4/M in · $1.6/M outRun now →
MiniMaxMiniMax

MiniMax M2

1M

The M2 release that put MiniMax on the map for cheap long-context agents.

$0.25/M in · $1/M outRun now →
MiniMaxMiniMax

Hailuo Text 01

1M

MiniMax's long-context text model from the Hailuo family.

$0.2/M in · $1.1/M outRun now →
NVIDIANVIDIA

Nemotron 3 Super

1M

NVIDIA's larger open MoE for teams building on the Nemotron stack.

$0.3/M in · $1/M outRun now →
NVIDIANVIDIA

Nemotron 70B Instruct

128K

NVIDIA's RLHF-tuned Llama derivative, a longtime arena favorite.

$0.12/M in · $0.3/M outRun now →
PerplexityPerplexity

Sonar

128K

Lightweight search-grounded model for fast, cited answers.

$1/M in · $1/M outRun now →
PerplexityPerplexity

Sonar Reasoning Pro

128K

Search-grounded reasoning with citations, for research-grade answers.

$2/M in · $8/M outRun now →
InceptionInception

Mercury 2

128K

The diffusion LLM that started it all — parallel token generation, low latency.

$0.25/M in · $1/M outRun now →
InceptionInception

Mercury Coder

128K

Diffusion-based coding model, one of the fastest ways to generate code.

$0.25/M in · $1/M outRun now →
ByteDance SeedByteDance Seed

Seed 1.8

262K

ByteDance's general-purpose Seed model with strong multilingual skills.

$0.3/M in · $1.2/M outRun now →
ByteDance SeedByteDance Seed

Doubao Pro 256K

256K

The model behind Doubao, China's most-used consumer AI assistant.

$0.4/M in · $1.5/M outRun now →
TencentTencent

Hunyuan 3 Pro

262K

Tencent's previous Hunyuan flagship for chat and enterprise copilots.

$0.5/M in · $2/M outRun now →
TencentTencent

Hunyuan Turbo

128K

Latency-optimized Hunyuan for interactive consumer workloads.

$0.2/M in · $0.8/M outRun now →
LiquidAILiquidAI

LFM2 8B

32K

Liquid's small liquid-architecture model for edge-class latency.

$0.02/M in · $0.05/M outRun now →
LiquidAILiquidAI

LFM2.5 40B

128K

Larger Liquid model for extraction, RAG, and structured outputs.

$0.1/M in · $0.25/M outRun now →
CohereCohere

Command A

256K

Cohere's enterprise flagship, tuned for RAG, tool use, and multilingual work.

$2.5/M in · $10/M outRun now →
CohereCohere

Command R+

128K

The RAG workhorse that made Cohere an enterprise staple.

$0.5/M in · $1.5/M outRun now →
CohereCohere

Aya Expanse 32B

128K

Open multilingual model covering 23 languages, from Cohere Labs.

$0.05/M in · $0.15/M outRun now →
AI21AI21

Jamba 1.6 Large

256K

AI21's hybrid SSM-transformer flagship with very fast long-context inference.

$2/M in · $8/M outRun now →
AI21AI21

Jamba 1.6 Mini

256K

Compact Jamba for long documents on a budget.

$0.2/M in · $0.4/M outRun now →
AmazonAmazon

Nova 2 Pro

1M

Amazon's frontier multimodal model, deeply integrated with AWS tooling.

$1.25/M in · $5/M outRun now →
AmazonAmazon

Nova 2 Lite

1M

Low-cost Nova for high-throughput enterprise pipelines.

$0.3/M in · $1.2/M outRun now →
AmazonAmazon

Nova Micro

128K

Text-only micro model, among the cheapest hosted options anywhere.

$0.035/M in · $0.14/M outRun now →
MicrosoftMicrosoft

Phi 5

128K

Microsoft's small-model line, trained on textbook-quality synthetic data.

$0.1/M in · $0.4/M outRun now →
MicrosoftMicrosoft

Phi 4 Multimodal

128K

Tiny multimodal Phi for vision and audio tasks at edge-friendly cost.

$0.07/M in · $0.3/M outRun now →
IBMIBM

Granite 4.0 H

128K

IBM's open enterprise model with hybrid Mamba-transformer architecture.

$0.08/M in · $0.3/M outRun now →
IBMIBM

Granite 3.3 8B

128K

Small open Granite for document tasks, popular in regulated industries.

$0.03/M in · $0.1/M outRun now →
falfal

Veo 3

video

Google's flagship video model with native audio generation.

$2.40 per clipRun now →
falfal

Kling 2.1 Master

video

Kling's top tier with the best motion quality in the lineup.

$0.80 per clipRun now →
falfal

Runway Gen-4

video

Runway's cinematic video model, a favorite for creative studios.

$0.50 per clipRun now →
falfal

Pika 2.2

video

Playful video model with strong stylization and fast turnaround.

$0.20 per clipRun now →
falfal

PixVerse V5

video

Budget-friendly video generation with surprisingly clean motion.

$0.15 per clipRun now →
falfal

Hunyuan Video

video

Tencent's open video model, a solid mid-price default.

$0.30 per clipRun now →
falfal

Wan 2.2

video

Alibaba's open video model — cheap, fast, and good enough for drafts.

$0.10 per clipRun now →