$ yogi.compute
Creator rewards, spent as compute. Hold 100,000 $YOGI and the API turns on. Sell, and it turns off. There is nothing to buy and no balance to top up.
Inference
Token
Every trade pays a fee
$YOGI generates creator fees on every trade. Those fees pay for inference.
Holding is the subscription
Hold 100,000 $YOGI and the API turns on. Sell, and it turns off.
Then use it, unmetered
Not a quota and not a share. Every holder gets the same access to all models.
Every model, unmetered
Prices are what each one costs the treasury, not what it costs you.
Claude Opus 5
Anthropic's flagship model for demanding reasoning, coding, and long-horizon agentic work.
GPT-6 Astra
OpenAI's flagship model for demanding end-to-end work: analysis, software engineering, deep research.
Claude Fable 5.1
Mythos-class model with the biggest gains in agentic coding and long-running workflows.
Grok 4.6
Frontier performance on coding, knowledge work, and STEM.
GLM 5.3
Large-scale reasoning model built for complex software engineering and long-horizon agent tasks.
Qwen3.8 2.4T A95B
Open-weight sparse mixture-of-experts model, the open-weight variant of Qwen3.8 Max.
Gemini 3.8 Flash
Google's most intelligent Flash model, with significant gains across coding and agentic tasks.
DeepSeek V4 Pro 0813
The GA release of DeepSeek V4 Pro.
Muse Spark 1.3
Multimodal reasoning model from Meta for long-running agentic, multi-agent, and coding workflows.
Kimi K3
2.8T parameter open-weight multimodal reasoning model for complex coding and knowledge work.
GPT-5.6 Sol
Flagship model in the GPT-5.6 series, suited for complex reasoning, coding, and agentic workflows.
Claude Sonnet 5
Anthropic's most capable Sonnet-class model, with frontier performance across coding and agents.
Gemini 3.1 Pro Preview
Google's frontier reasoning model, delivering enhanced software engineering performance.
GPT-5.4
OpenAI's latest frontier model with a 1M+ token context window.
DeepSeek V4 Flash
Efficiency-optimized Mixture-of-Experts model with 284B total and 13B activated parameters.
Qwen3.8 Flash
Multimodal reasoning model for coding assistance, agentic workflows, and visual understanding.
Mistral Medium 3.5
Dense 128B instruction-following model supporting text and image inputs.
MiniMax M3
Multimodal foundation model with text, image, and video inputs and a 1M-token context.
Grok 4.20
Reasoning model with industry-leading speed and agentic tool calling capabilities.
Nemotron 3.5 Lightning
Open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total.
Sonar Pro
Perplexity's advanced search-grounded model. Pricing includes search pricing.
Mercury 2.5
The fastest reasoning LLM and the latest diffusion LLM, generating tokens in parallel.
Seed-2.0-Code
Model optimized for agentic coding, frontend development, and multilingual programming.
Hy4 preview
Mixture-of-experts model with 49B active parameters out of 770B total, designed for coding agents.
GPT-5.4 Mini
Core GPT-5.4 capabilities in a faster, more efficient model for high-throughput workloads.
Claude Haiku 4.5
Anthropic's fastest and most efficient model, delivering near-frontier intelligence.
Gemini 3.5 Flash Lite
High-efficiency model with upgraded agentic capabilities for focused subagent workloads.
Llama 4 Maverick
High-capacity multimodal model from Meta, built on a mixture-of-experts architecture.
Qwen3.6 27B
Dense 27-billion-parameter language model with hybrid multimodal capabilities.
GLM 4.7 Flash
30B-class SOTA model balancing performance and efficiency, optimized for agentic coding.
Veo 3 Fast
Google's Veo 3, fast tier, with audio. The most expensive request this API serves.
Kling 2.5 Turbo Pro
Kling's fast tier. Reliable prompt adherence and clean motion.
Hailuo 02
MiniMax's video model. Good physical motion and camera movement.
LTX Video 13B
The cheapest way to get a clip out of this API. The right default for iterating on a prompt.
LFM2.5-2.6B (free)
Compact reasoning model suited for agent workflows, data extraction, and RAG.
Gemma 4 31B (free)
Google DeepMind's 30.7B dense multimodal model with a 256K token context window.
o5
OpenAI's dedicated reasoning model for hard math, science, and multi-step logic.
o4 Mini
Compact reasoning model with strong STEM performance at a fraction of o5's cost.
GPT-5.4 Nano
The cheapest GPT-5.4 tier, built for high-volume classification and extraction.
GPT-5.2
Previous-generation frontier model, still a strong default for general assistant work.
GPT-5.1 Codex
Coding-specialized GPT-5.1 variant tuned for agentic software engineering.
GPT-4.1
Reliable 1M-context workhorse from the GPT-4 series, cheap for long-document tasks.
Claude Opus 4.8
Previous Opus generation with deep reasoning and careful long-form writing.
Claude Sonnet 4.6
Balanced Sonnet release known for dependable coding and instruction following.
Claude Haiku 4
Fast, inexpensive Haiku for latency-sensitive assistants and routing.
Gemini 3.5 Pro
Google's Pro-tier reasoning model with strong multimodal and coding scores.
Gemini 3 Flash
Fast Gemini for everyday chat, summarization, and light agentic work.
Gemma 4 12B
Small open Gemma model for on-device-class workloads served from the API.
Gemma 3 27B
Older open-weight Gemma, still popular for fine-tuned derivatives.
Grok 4
The original Grok 4 flagship with strong reasoning and real-time knowledge.
Grok 4.1 Fast
Speed-optimized Grok with a 2M context, built for agentic tool calling at scale.
Grok Code Fast 2
Coding-tuned Grok variant for autocomplete-style and edit-heavy workflows.
Llama 4 Scout
Lightweight Llama 4 with a huge context window, ideal for retrieval-heavy apps.
Llama 3.3 70B
The classic open-weight 70B that still powers half the fine-tunes in the wild.
Muse Spark 1.2
Earlier Muse Spark release with solid multimodal reasoning at a lower price.
DeepSeek R2
DeepSeek's next-gen reasoning model with long chains of thought and tool use.
DeepSeek V3.2
The V3 refresh that made DeepSeek famous: strong chat quality, tiny price.
DeepSeek R1
The original open reasoning model that started the RL-reasoning wave.
Qwen3.8 Coder
Code-specialized Qwen3.8 with repo-scale context and agentic editing skills.
Qwen3.6 72B
Dense 72B Qwen for teams that want open weights without MoE complexity.
Qwen3 VL 235B
Vision-language flagship of the Qwen3 line, strong on documents and UI screenshots.
QwQ 32B
Compact open reasoning model that punches far above its size on math.
Mistral Large 3
Mistral's flagship dense model for enterprise assistants and RAG.
Codestral 26
Mistral's code model, tuned for fill-in-the-middle and completion workloads.
Magistral Medium
Mistral's reasoning line with transparent, inspectable chains of thought.
Devstral 2
Agentic coding model built with SWE-bench-style tool use in mind.
Mistral Small 4
Small, cheap, and surprisingly capable Mistral for everyday tasks.
Kimi K2.5
The K2 refresh that made Kimi a favorite for long-context agentic coding.
Kimi K3 Thinking
Deliberate-reasoning variant of K3 for the hardest problems, no time pressure.
GLM 5
Z.ai's previous flagship, still strong on coding and agentic benchmarks.
GLM 4.7
Dependable mid-tier GLM for assistants, extraction, and routing.
MiniMax M2
The M2 release that put MiniMax on the map for cheap long-context agents.
Hailuo Text 01
MiniMax's long-context text model from the Hailuo family.
Nemotron 3 Super
NVIDIA's larger open MoE for teams building on the Nemotron stack.
Nemotron 70B Instruct
NVIDIA's RLHF-tuned Llama derivative, a longtime arena favorite.
Sonar
Lightweight search-grounded model for fast, cited answers.
Sonar Reasoning Pro
Search-grounded reasoning with citations, for research-grade answers.
Mercury 2
The diffusion LLM that started it all — parallel token generation, low latency.
Mercury Coder
Diffusion-based coding model, one of the fastest ways to generate code.
Seed 1.8
ByteDance's general-purpose Seed model with strong multilingual skills.
Doubao Pro 256K
The model behind Doubao, China's most-used consumer AI assistant.
Hunyuan 3 Pro
Tencent's previous Hunyuan flagship for chat and enterprise copilots.
Hunyuan Turbo
Latency-optimized Hunyuan for interactive consumer workloads.
LFM2 8B
Liquid's small liquid-architecture model for edge-class latency.
LFM2.5 40B
Larger Liquid model for extraction, RAG, and structured outputs.
Command A
Cohere's enterprise flagship, tuned for RAG, tool use, and multilingual work.
Command R+
The RAG workhorse that made Cohere an enterprise staple.
Aya Expanse 32B
Open multilingual model covering 23 languages, from Cohere Labs.
Jamba 1.6 Large
AI21's hybrid SSM-transformer flagship with very fast long-context inference.
Jamba 1.6 Mini
Compact Jamba for long documents on a budget.
Nova 2 Pro
Amazon's frontier multimodal model, deeply integrated with AWS tooling.
Nova 2 Lite
Low-cost Nova for high-throughput enterprise pipelines.
Nova Micro
Text-only micro model, among the cheapest hosted options anywhere.
Phi 5
Microsoft's small-model line, trained on textbook-quality synthetic data.
Phi 4 Multimodal
Tiny multimodal Phi for vision and audio tasks at edge-friendly cost.
Granite 4.0 H
IBM's open enterprise model with hybrid Mamba-transformer architecture.
Granite 3.3 8B
Small open Granite for document tasks, popular in regulated industries.
Veo 3
Google's flagship video model with native audio generation.
Kling 2.1 Master
Kling's top tier with the best motion quality in the lineup.
Runway Gen-4
Runway's cinematic video model, a favorite for creative studios.
Pika 2.2
Playful video model with strong stylization and fast turnaround.
PixVerse V5
Budget-friendly video generation with surprisingly clean motion.
Hunyuan Video
Tencent's open video model, a solid mid-price default.
Wan 2.2
Alibaba's open video model — cheap, fast, and good enough for drafts.