AI Inference Cost Structure
E249|Token经济转点:OpenClaw、Hermes到本地自研的Agent进化之路 adds a practitioner-level token-economics turn. 东旭 / Dongxu contrasts earlier high-spend Token Maxxing with a newer Token Efficient Agent Workflow in which frontier models, local models, cheaper open models, multi-agent review, and deterministic tools are routed by task value. 张宏江 adds the macro version: if token cost keeps falling, total use may still rise through Jevons Paradox In AI because agents create more tasks and longer loops.
More Trillion Dollar IPOs, Anthropic $3T, Zuck’s Price War, China Ends Open Source?, Trump Accounts adds an enterprise-budget shock. Chamath Palihapitiya describes model spend doubling and says routing through OpenRouter, GLM 5.2, and cheaper providers produced roughly 95% savings; the source turns token cost into a board-level Enterprise AI ROI Audit issue rather than only a developer quota problem.
Open Source Wins, AGI Is Here, and Scorsese’s AI Toolkit with CEOs of Cerebras & Black Forest Labs adds Andrew Feldman’s reasoning-as-inference version through Cerebras. Feldman says reasoning models spend many internal tokens, so faster inference changes not only cost but elapsed time, guardrail latency, and whether 24- to 48-hour reasoning loops become usable workflows.
Bill Maris: How Google Could Crush AI Competitors, Why Small Funds Win, and AI’s Atari Stage adds the strategic price-war version through Bill Maris. Maris argues that Google could cut Gemini token costs sharply enough to pressure OpenAI and Anthropic, making Google AI Token Price Leverage a case where inference cost becomes a competitive weapon rather than only a developer budget constraint.
Microsoft CEO Satya Nadella on AI’s Business Revolution: What Happens to SaaS, OpenAI, and Microsoft? | LIVE from Davos adds Azure as a hyperscaler “token factory” case through Satya Nadella. The source pushes inference cost beyond API prices into heterogeneous hardware, utilization, total cost of ownership, model routing, and the ability to serve agent workloads through Microsoft Foundry.
巴黎水和圣培露还能赚钱,雀巢为何要剥离水业务? adds another pricing signal through DeepSeek and Qwen. The source says DeepSeek planned a substantial API price increase after using peak/off-peak pricing, while Alibaba’s next Qwen could remain open source but seek revenue sharing from large customers. This reinforces that model access, compute peaks, open-model distribution, and monetization cannot be separated indefinitely.
「蜘蛛侠」新片拿下近半国内票房,AI 模型爆发价格战 adds a daily-news price-war example. The episode says Qwen’s new flagship model, Kimi K3, a source-named Deepsea product, and OpenAI price cuts are all part of a market where falling model prices and narrowing capability gaps make task-level model choice more important.
Featherless AI: When Your Weekend Experiment Makes More Than Your Startup adds Featherless AI’s long-tail hosted-inference case. Eugene Chia says many providers avoid rarely used models because keeping GPUs warm for each model is uneconomic, while GPU Hot Swapping can bring requested models online quickly and let one GPU serve different customers and models dynamically. The source also adds Flat-Rate AI Inference Pricing as a customer-facing response to complex per-model and usage-based pricing.
177: 详解Kimi K3:强到冲击Anthropic估值的模型什么样? adds a detailed Kimi K3 serving case. Zhao Chenyang / 赵晨阳 decomposes inference experience into first-token latency, single-user decode speed, and queue time, then argues that agent tasks should be priced by total task completion rather than only by dollars per million tokens. The source also adds an architecture-specific cost issue: KDA can lower long-context memory movement but complicates Prefix Caching because recurrent state is mutable.
E246|何谓蒸馏?聊聊硅谷如何看中国开放模型逼近前沿 adds the open-model price-pressure version. 王铁镇 splits closed model API pricing into hardware inference cost plus closed-model premium, then argues that strong open-weight models let inference providers compete closer to hardware and service cost. The same source adds Agent Inference Workload: long inputs, short outputs, prefix/KV-cache reuse, scheduling, and hardware/software co-design can decide whether agent-heavy serving is economical.
270.大厂押注AI办公,飞书和钉钉却先成了配角 adds the consumer-scale liability version through Doubao. The source asks whether very large C-end DAU is an asset or liability when every chat can generate token cost, and argues that AI office and enterprise workflows may create a stronger payment path than undifferentiated chatbot traffic.
148. 对游凯超3小时访谈:开源Infra、和模型Co-design 、“如果vLLM失败,我们会后悔一辈子” adds the inference-engine view through vLLM. 游凯超 argues that serving cost is shaped by Continuous Batching, attention-state management, Prefix Caching, model-architecture support, and hardware fit, not only by visible token volume or published API prices.
AI inference cost structure is the idea that large-model services incur meaningful cost every time users generate output. In 从QQ会员到豆包包月,中国人为什么总觉得软件该免费, the hosts contrast this with older internet services where serving more users often had much lower marginal cost after the platform was built. Eric Ries on How Founders Quietly Lose Their Company adds a SaaS-founder version: AI can move software away from near-zero marginal cost assumptions because advanced model usage consumes tokens and production infrastructure. Agent 元年第 500 天:什么在消失,什么在诞生——为什么我们不该再投资 GUI 思维的软件? adds that token availability and effective quota can fluctuate with energy, model capability, demand, and infrastructure supply. 对话 MiniMax 闫俊杰:M3、10X 计划、10T 模型、和智能的终局 adds a developer-workflow version: MultiCard watches token spending across tools such as Claude Code, Codex, and Cursor, while MiniMax uses very large token usage as one signal of model adoption.
EP108 Vibe Coding大地震:Cursor定价争议、Windsurf收购风波,模型厂商亲儿子们又将如何进场? adds the Vibe Coding pricing case. Cursor’s pricing controversy is framed as the moment when long-context coding, background agents, bug agents, and heavy interactive use made the old flat request-count model hard to sustain.
Vol. 161 从开发自己的 OpenClaw 聊起 adds an always-on personal-agent case. The hosts describe Open Claw-style polling, many AI Skills, web coding, and persistent triggers as token-heavy patterns, and they doubt easy commercialization when the agent must keep spending model calls to stay useful.
Vol. 166 闲聊: 从 Gemini 到 AI 的加速与混沌 adds the user-psychology side of token cost. The hosts connect expensive frontier APIs, subscription tiers, quota resets, and review-heavy coding workflows to AI anxiety and the feeling that paid capacity must be used.
EP124 为什么 Agent 时代,CLI 反而成了最优解?⚡ adds a content-consumption and CLI-design case through Podwise. Skills can make users and agents process far more podcast, documentation, and product material, increasing token flow; at the same time, an Agent-Optimized CLI can reduce waste by moving stable exports and format conversions into deterministic local commands instead of asking the model to regenerate them.
EP101 对话 Simon:AI 创业者的第一项基本功是把账算明白 adds an AI companion and game-social case through Simon and Mico AI Lab. The episode argues that Character AI-style chat can become more expensive as relationship history deepens, because better experience requires memory retrieval and longer prompts, while the paying segment may not be large enough to absorb that cost.
20 个问题,搞懂 OpenClaw:爆红机制、本质变化、创业机会 adds a user-behavior case through Open Claw. 鸭哥 notes that expensive model calls can make users hesitate before delegating complex work, while cheaper or subscription-style access encourages experimentation. Long-running agents also spend tokens on memory, context compaction, tool use, and repeated feedback loops.
「1 亿 Token 俱乐部」挤爆了,AI 的燃料不够了:对谈于文渊 adds the cloud-serving side through Aliyun Bailian. Yu Wenyuan argues that token counts are not comparable unless model type, intelligence, latency, throughput, peak demand, GPU scheduling, and stability are considered. This turns inference cost into MaaS Infrastructure: a platform must keep compute utilized while still delivering secure, fast, reliable tokens.
Vol. 170 Fable 5 重出江湖,GPT 仍需努力 adds a heavy-user workflow case through Fable 5. The hosts describe Fable-specific limits, faster quota burn than previous models, an API change costing about five dollars, and the danger of running full Superpowers flows on expensive models. The episode’s practical response is Model Routing Cost Control: route simple work to cheaper models and reserve top models for planning, hard coding, and review judgment.
把 AI 吹成核武器的人,亲手拉下了新冷战铁幕 adds a domestic open-model serving case through GLM 5.2. The hosts highlight long context and improved coding behavior while noting slow speed, which they interpret as possibly related to compute constraints. The source also links cost and capacity to SaaS Reliability Under Policy Risk: customers may value models they can access reliably even if they are not the absolute strongest.
Vol. 167 Token 如流水,Agent 似朝阳 adds the day-to-day heavy-user case. The hosts say unconstrained Codex and Claude Code usage would be very expensive at API prices, mention companies limiting employee AI API access, and connect token cost to broader infrastructure substitution such as cheaper models, local models, deterministic scripts, and Cloudflare services.
Vol. 162 科技快乐星球44: 新模型“SOTA们”齐贺新春 adds a release-cycle and infrastructure layer. The hosts test Gemini model-version costs in their own tool, discuss ChatGPT Go and possible ads as price segmentation, and connect cloud-chip commitments, power, data centers, and speculative space compute to the cost of serving frontier models.
E155.似乎没什么人再提「AI 泡沫论」了 adds Jevons Paradox In AI as the demand response to falling token cost. The source argues that cheaper tokens can increase total consumption because users run more rounds, agents execute more steps, and more workflows move into production. It therefore treats token growth as both an adoption signal and a cost/infrastructure stress signal inside AI Investment Metrics.
存储三巨头破万亿市值,存储超级周期何时能见顶?| S10E13 adds the memory-cost version. The source argues that inference can be more memory-intensive than training because long contexts and KV cache require fast memory, and that long agent workflows also create Agent-Era NAND Storage demand for recoverable intermediate state.
141. Freda的投资札记第2集:Tokenmaxxing、把电机塞进蒸汽机、接力赛变篮球赛、孤独、人的连接 adds the investor-facing Token Maxxing correction. Freda / Friday argues that total token consumption must be decomposed into users, tasks per user, token per task, and dollar per token. Agent workflows can raise total consumption while model and workflow optimization reduce waste, so raw token growth is neither pure value nor pure bubble by itself.
263.Sora死了,Adobe跌了,美图何去何从? adds a creative-tool and AI-video case. The source frames Sora as struggling partly because video inference cost and product revenue did not match, and frames Adobe as pressured because AI features embedded in professional tools can create costs before they create enough paid uplift. Meitu / 美图’s Model Container Strategy is presented as one way for an application company to avoid carrying the full foundation-model cost burden.
EP270 一枚芯片的漫长征途:我们离“算力自由”还有多远? adds the semiconductor supply side of the same cost structure. The episode argues that token prices depend on chips, GPUs, manufacturing, packaging, power, software ecosystems, and MaaS Infrastructure, then compares future token-price declines to mobile data getting cheap enough to enable new app behavior.
E230|1万亿收入预期背后:英伟达的巅峰与软肋 adds Inference as Cash Flow and Token per Watt. The episode argues that production inference can become recurring cash flow for infrastructure suppliers, but only if Nvidia systems, High Bandwidth Memory, interconnect, power, cooling, and GPU Cloud Operations convert hardware into reliable tokens at acceptable cost.
E228|谷歌TPU能撼动英伟达吗?前TPU工程师首次揭秘 adds the TPU route to lower inference cost. Henry argues that Google can reduce TCO for large-scale serving when High-Throughput Inference Batching, TPU Pods, XLA, High Bandwidth Memory, and data-center deployment are aligned. The source also sets a boundary: single-user, low-latency agent calls may fit Groq-style low-latency chips better than TPU’s high-throughput optimization.
快一点!再快一点!快到世界能实时生成|和生数科技张金涛聊:Vidu S1、推理加速、实时交互视频 adds the real-time video version. Vidu S1 makes inference cost visible as continuous frame generation: every live session consumes model capacity, and the source says the API was priced around two to three yuan per minute while web and app access were free at that stage. Acceleration therefore shapes both product feasibility and unit economics.
OpenClaw 之后,我只想未来 3-6 个月的事情|对谈 Sheet0 创始人王文锋 adds a small-team management case through Sheet0. 王文锋 / Wang Wenfeng says the team spent about $20,000 on AI coding in the prior month and treats high token availability as a condition for training people to use agents intensely. The source also suggests that customers may be segmented by token consumption or replaced labor budget rather than by company headcount alone.
用 Agent 动力学,和 40 个 Agents 一起为「人 + AI」做产品|对谈 Slock.ai 创始人 RC adds a high-token multi-agent team case through Slock.ai. A seven-person team coordinating roughly forty agents turns inference budget into organization infrastructure: token spend supports parallel work, shared memory, model diversity, and agent review, but only matters if it converts into accepted product output.
我们是如何定义 OpenClaw for Teams 新产品形态的|对谈 Kuse&Junior 联创兼 CTO 宇豪 adds Kuse’s product-pricing and internal-agent case. Yuhao / 宇豪 says fixed pricing failed for agentic work because users could unknowingly trigger many rounds of expensive model calls; internally, Kuse’s long-running agents cost more than $20,000 per month in tokens, making the comparison against hiring and organization friction explicit.
AI 不只比智商,WAIC 和 Kimi K3 透露了什么新竞争 adds two concrete cost cases. The Kimi K3 podcast-agent build reportedly took about three hours and roughly 250,000 tokens, showing how model choice affects latency and token-per-task even when the output is useful. The same source’s Speech To Text Cost Optimization case reports reducing one hour of audio transcription from about 0.6 yuan to under 0.1 yuan through engineering optimization and batch processing rather than only model fine-tuning.
「模型能力已经够了,要卷就卷 infra」|对谈戴冠兰:Runta 创始人 adds Runta’s runtime-cost view. 戴冠兰 says customers now ask about budgets, token use, compute spending, where agents run, and how they scale; Runta also analyzes whether a harness wastes tokens and where that waste appears.
Vol. 171 假如我们有无限 Token adds the heavy-subscriber and true-abundance contrast. The hosts describe a spectrum from fast-mode quotas and multiple paid accounts to API/OpenRouter spending and the idea of genuinely unlimited token access. The source’s cost contribution is behavioral: once users believe tokens are plentiful, they start more background tasks and long loops, so the cost question moves from “can I afford this call?” to “does this growing queue become accepted work?”
Vol. 172 Codex 卖重置套餐,DeepSeek 峰谷调价,苹果重回 5 万亿等 adds two explicit user-visible price mechanisms. First, possible Codex reset purchases turn subscription quota recovery into an add-on product. Second, DeepSeek peak/off-peak pricing turns serving load into Peak-Valley AI Inference Pricing, where users can move batchable work to cheaper windows but cannot always reschedule interactive text, coding, or assistant tasks.
Key Claims
- Token generation, GPU capacity, electricity, storage, and infrastructure procurement make AI usage costly at scale.
- Free growth is harder when user growth directly increases inference load.
- The cost applies to large products such as Doubao and smaller products such as 你的书房.
- AI pricing therefore becomes a product and infrastructure problem, not just a marketing decision.
- AI-assisted SaaS prototypes need unit economics checks before founders assume a demo can become a profitable production product.
- Agent economies depend on token costs becoming low and reliable enough for long-running work.
- Cost-aware teams may orchestrate multiple models instead of assigning every task to the largest or most expensive model.
- High token usage can signal adoption, but it also increases pressure to improve serving efficiency and workflow value.
- AI coding tools may need pricing that maps closer to real token/API consumption, but users still need understandable remaining-budget signals.
- Always-on personal agents need trigger discipline because periodic checks and open-ended skill use can turn small automations into ongoing inference spend.
- Subscription plans and quota resets can change behavior even before direct API bills arrive, because users may work around perceived scarcity or try to exhaust paid capacity.
- Skills can increase both useful content consumption and inference cost; product design needs quota visibility, stable local tooling, and judgment about whether a task is work-value or entertainment-value.
- In companion-chat products, relationship depth can increase inference cost because useful memory and context grow with use.
- In executable agents, cost affects delegation psychology: users may underuse capable agents if every long task feels like a billable risk.
- Raw token counts can mislead because different models and workloads consume very different compute for the same visible token volume.
- Serving-side economics include latency, peak capacity, scheduling, security, GPU utilization, and domestic compute supply, not just per-token API price.
- Cost-aware model routing becomes a user workflow problem, not only a cloud infrastructure problem, when top models have separate limits and burn noticeably faster.
- Long-context open models still have serving constraints; speed, capacity, and access reliability shape whether they can substitute for restricted closed models.
- Heavy personal and enterprise agent use makes cost visible even before a direct bill arrives, because API prices, subscription limits, task decomposition, and alternative services change actual workflow choices.
- Published model prices are not enough for users: actual cost depends on the task, version behavior, latency, quota, and how much verification or repair a model induces.
- Lower per-token cost can increase total token demand when agents and applications expand the number of calls.
- Token-per-task matters because visible output length, hidden reasoning tokens, model retries, and verification work can differ sharply across models completing the same task.
- Inference engines can lower cost by improving scheduling, cache reuse, and model/hardware fit; prompt and harness design can raise cost if it breaks otherwise reusable context.
- Creative-media AI can make cost visible faster than older software because video generation, image iteration, quality control, and retries all consume model capacity before the user is clearly willing to pay more.
- Agent-runtime cost includes more than model price: scheduling, sandbox duration, scaling, token waste, retries, GPU timing, and execution recovery can all determine whether an agent task is economical.
- Inference cost includes memory bandwidth, memory capacity, NAND checkpointing, and storage hierarchy choices, not only GPU arithmetic or model-token pricing.
- Token prices can become an application-shaping constraint: cheaper compute may make new AI products viable, while high cost splits advanced and free-model experiences.
- Inference demand can behave like recurring cash flow, but the margin depends on energy efficiency, memory movement, utilization, and operational reliability.
- Token-per-watt makes energy and data-center constraints visible inside model-serving economics.
- For agent-heavy teams, token spend becomes a management variable: the relevant comparison is not only API cost, but whether the spend converts into accepted work, fewer hires, faster review, or better product output.
- Multi-agent collaboration can multiply inference demand because coordination, progress summaries, memory updates, reviews, and parallel attempts all consume tokens alongside final output.
- Agentic products may need usage-based, credit-based, or salary-like pricing because both customer tasks and internal AI employees create variable inference exposure.
- Specialized accelerators can lower inference cost only when workload shape, batching, compiler support, memory bandwidth, and utilization are favorable; otherwise cheap nominal hardware can still produce expensive tokens.
- Real-time video generation adds a per-minute, frame-rate-sensitive cost pattern where speed, quality, and live-session duration determine whether a product can be served economically.
- Token-per-task can differ sharply across models: a slower cheaper model may still be useful for background work, while an optimized pipeline may beat a stronger generic model for repeated audio transcription.
- Open-weight models can compress closed API premiums by letting hosted inference providers compete on hardware cost, scheduling, cache reuse, and service quality.
- Agent workloads can make KV-cache lifetime, prefix reuse, and hardware/software co-design as important as published token prices.
- At consumer scale, high DAU can become a cost liability unless the product can convert repeated token use into subscriptions, ads that preserve trust, commerce, or enterprise workflow revenue.
- Long-context architecture can shift costs from simple KV cache growth into recurrent-state lifecycle, rollback, kernel design, and prefix-reuse engineering.
- Dynamic model loading can make Long-Tail Model Hosting viable when low-volume models would otherwise waste standby GPU capacity.
- Flat-rate pricing reduces user anxiety around AI bills, but it depends on internal utilization, limits, and cost controls remaining sound.
- Model price cuts can expand usage while also compressing margins, so providers need either better utilization, differentiated workflows, or higher total demand to offset lower per-call revenue.
- Vol. 171 adds that perceived abundance can increase total task creation even before real marginal cost disappears; users may spend subscription capacity on research, tests, reviews, migrations, and generated tools that then require human acceptance work.
- Vol. 172 adds that AI cost is now temporal as well as volumetric: reset timing, peak windows, random task arrival, and model-specific quality all change the effective price of useful work.
- Nadella’s All-In source adds that infrastructure providers compete on turning heterogeneous compute into reliable tokens with good utilization, not only on headline model access.
- Maris’s All-In source adds that a full-stack provider can use token-price cuts to win installed base and compress standalone model-company economics.
- Feldman’s All-In source adds that reasoning quality, latency, guardrails, and recursive loops are coupled: spending more inference can improve answers, but only if the system can deliver the work fast enough and cheaply enough to matter.
- E249 adds that expert users may reduce direct bills by shifting repetitive work to local models, but the agent market can still grow because cheaper loops invite more tasks into the workflow.
Connections
- AI Commercialization Pressure — broader business pressure created by high model costs.
- AI Subscription Economics — pricing model used to manage ongoing usage costs.
- ByteDance and Doubao — large-scale case in the source.
- Data Portability And Sustainable Tools — design response for smaller products that want lower operating burden.
- Validated Learning — product experiments should test economics as well as customer interest.
- Agentic Economy and Token Grant — agent-era examples where token supply becomes an input to creation.
- MiniMax M3 and MultiCard — coding-workflow case where token cost influences model orchestration.
- Frontier Model Scaling — related training-side pressure around model size, data, and compute.
- Cursor, Vibe Coding, and Model Provider Tool Competition — AI coding case where token cost reshapes product strategy.
- Open Claw, On-Demand Apps, and Agent Native Software — personal-agent case where token cost limits product viability.
- Peak-Valley AI Inference Pricing, Codex, DeepSeek, AI Subscription Economics, and Model Routing Cost Control — Vol. 172’s reset and peak/off-peak pricing branch.
- Codex, Claude Code, Superpowers, and AI Workforce Monitoring — personal workflow and measurement cases added by Vol. 166.
- Podwise, Agent-Optimized CLI, and AI Skills — content-processing and deterministic-tooling case added by EP124.
- AI Startup Unit Economics, Character AI, and Mico AI Lab — AI game/social commercialization case added by EP101.
- Open Claw, IM Agent Interfaces, Local Agent Execution, and Agent Harness — long-running agent cost case added by the 20-question source.
- Aliyun Bailian, Yu Wenyuan, and MaaS Infrastructure — serving-platform case where compute-to-token conversion becomes the main infrastructure problem.
- Fable 5, Superpowers, GrillMe Skills, and Model Routing Cost Control — heavy-user workflow and quota-control case added by Vol. 170.
- GLM 5.2, Open Source AI Models, and SaaS Reliability Under Policy Risk — long-context, speed, and access-reliability case added by the Keji Luandun export-control episode.
- Codex, Claude Code, DeepSeek, Cloudflare, and Model Routing Cost Control — heavy-user cost and substitution case added by Vol. 167.
- Gemini, Amazon, Anthropic, MaaS Infrastructure, and AI Subscription Economics — model-version cost testing, cloud-chip binding, and pricing case added by Vol. 162.
- Jevons Paradox In AI, AI Investment Metrics, and Human Resource Deflation Compute Infrastructure Inflation — E155’s efficiency-to-demand and investment-metric extension.
- Token Maxxing, Freda / Friday, and AI Investment Metrics — episode 141’s decomposition of token growth into users, tasks, token efficiency, and price.
- Sora, Adobe, Meitu / 美图, and Model Container Strategy — creative-media and application-company cost case added by Luanfanshu.
- Memory Wall, High Bandwidth Memory, Agent-Era NAND Storage, and AI Data Center Memory Hierarchy — memory-cost branch added by What’s Next.
- Compute Freedom / 算力自由, Domestic AI Chip Catch-Up, Semiconductor Supply Chain, and AI Compute Continuity — semiconductor and cost-availability branch added by EP270.
- Inference as Cash Flow, Token per Watt, Nvidia Blackwell Platform, Nvidia Vera Rubin Platform, GPU Cloud Operations, and Data Center Power Bottleneck - recurring inference and system-deployment branch added by E230.
- Sheet0, 王文锋 / Wang Wenfeng, AI Managing AI, and One-Person Company — small-team token-budget and labor-substitution frame added by the 42章经 source.
- Slock.ai, RC, Agent Dynamics, and Model Workflow Fit — multi-agent token-budget and model-diversity case added by the RC episode.
- Kuse, Junior, Outcome-Based AI Pricing, and AI Organization Design — fixed-pricing failure and internal-agent token-budget case added by the Yuhao source.
- TPU, High-Throughput Inference Batching, XLA Compiler, TPU Pod System Optimization, Ironwood TPU, and Low-Latency Inference Chip - E228’s TPU inference-cost boundary.
- Vidu S1, Streaming Video Generation, SAGE Attention, TurboDiffusion, and Inference Acceleration Stack — real-time video inference-cost case added by the Shizilukou Crossing source.
- Kimi K3, Speech To Text Cost Optimization, Model Workflow Fit, and Top Model Build Runtime Split — token-per-task and optimized audio pipeline cases added by Keji Luandun.
- OpenRouter, Neo Cloud, Closed Model API Moat Pressure, Open-Weight Commercial Licensing, and Agent Inference Workload — open-model hosted-inference and agent-serving branch added by E246.
- Doubao, Feishu / 飞书, Doubao Enterprise Edition / 豆包企业版, AI Office Agent, and AI Commercialization Pressure - consumer-scale cost and office-work monetization branch added by Luanfanshu episode 270.
- Kimi K3, Kimi Delta Attention / KDA, Prefix Caching, Kernel Development Agents, and Agent Inference Workload - long-context serving branch added by LateTalk episode 177.
- Featherless AI, Eugene Chia, GPU Hot Swapping, Long-Tail Model Hosting, and Flat-Rate AI Inference Pricing - dynamic hosted-inference and pricing branch added by The SaaS Podcast.
- Qwen, Kimi K3, OpenAI, Anthropic, and Model Routing Cost Control - price-war and task-fit branch added by 声动早咖啡.
- Unlimited Token Workflow, AI Subscription Economics, OpenRouter, Token Maxxing, and AI Use Pacing - abundant-token behavior and review-debt branch added by Vol. 171.
- Azure, Token Factory AI Infrastructure, Microsoft Foundry, and AI Model Orchestration - Microsoft cloud-serving branch added by All-In.
- Google AI Token Price Leverage, Google, Gemini, OpenAI, and Anthropic - Maris interview branch around token price as platform leverage.
- Andrew Feldman, Cerebras, Low-Latency Inference Chip, Token Maxxing, Loop Maxxing, and Frontier Model Release Governance - All-In reasoning-inference and guardrail-latency branch.
- 东旭 / Dongxu, 张宏江 / Zhang Hongjiang, Token Efficient Agent Workflow, Local Agent Execution, and Multi-Agent Collaboration — E249’s routing, local-cost, and agent-loop extension.