Updated · 32 episodes · 12 shows · 32 source notes
DeepSeek
Overview
DeepSeek is a Chinese AI lab and open-weight model family that became a major reference point for reasoning models, constrained-compute engineering, model-infrastructure co-design, low-cost inference, and Chinese AI competition.
Current Profile
Across the bounded sources, DeepSeek is significant less as a single cheap-model headline than as a combined technical, ecosystem, and market event. R1 and V3-era releases made architecture, sparse models, reinforcement learning, distillation, open releases, inference optimization, and constrained-resource engineering visible to a broad audience. They also forced investors, model labs, infrastructure projects, Chinese technology companies, and enterprise buyers to reconsider whether access to the largest compute budget creates a durable moat by itself.
The current judgment remains qualified. Reported training costs do not capture the full model factory, open weights do not expose every dataset or process, and public distillation accusations remain unproven in the source set. DeepSeek’s later pricing, local-use, routing, and harness examples show a normalizing provider: it must fund inference, compete on quality and reliability, support infrastructure, and coexist with other models rather than remain a one-time geopolitical symbol.
Key Characteristics
- Efficiency-oriented model engineering can change capability per unit of compute and weaken a simple spending-only moat.
- Open releases spread technical learning, downstream deployment, inference-engine work, and ecosystem adoption beyond the original lab.
- DeepSeek is both a technical actor and a market/geopolitical signal affecting China-tech narratives, U.S. AI valuation, and export-control debates.
- Real-world use increasingly depends on routing among hosted, local, cheap, and frontier models rather than choosing one permanent provider.
- The model family’s influence includes architecture, post-training, distillation practice, and model-infrastructure co-design, not only price.
- Enterprise deployment retains trust, provenance, jurisdiction, safety, and data-control constraints even when weights can be self-hosted.
Evidence
Efficiency, architecture, and model-infrastructure co-design
- Vol.114 AI的2025和DeepSeek们的未来 | 对谈复旦张奇教授 treats DeepSeek as an engineering-efficiency and MoE cost-structure shock while preserving post-training and statistical-model limits.
- 从蒸馏到合成数据到 RSI,模型竞争的下一个焦点是什么?|对谈 Evolvent AI 联创孟繁青 uses Multi Latent Attention to argue that architecture and constrained-resource innovation matter alongside distillation.
- 148. 对游凯超3小时访谈:开源Infra、和模型Co-design 、“如果vLLM失败,我们会后悔一辈子” presents DeepSeek as a major vLLM catalyst and a strong Model-Infra Co-Design case.
- EP 34: DeepSeek R1 vs GPT-4: The $6M Model That Changed AI Economics records the broad-industry efficiency shock while keeping the reported cost and benchmark comparison source-scoped.
Open ecosystem and technical diffusion
- E246|何谓蒸馏?聊聊硅谷如何看中国开放模型逼近前沿 uses R1-style openness to distinguish legitimate reuse from unsupported copying claims and to connect model progress with Scaling Efficiency.
- 174. 我们还能给算法当多久的品味老师?|对谈亚马逊AGI查晟 says open releases spread architectural learning while weakening the originating lab’s exclusive data flywheel.
- 179: 蒸馏风暴:一场无人公开谈论的技术竞赛 documents R1’s distilled-model influence and separates public suspicion from complete evidence.
- Vans、匡威风光不再,经典帆布鞋为什么卖不动了? adds the later DeepSeek Harness developer-preview branch.
Cost, routing, and local execution
- Vol. 172 Codex 卖重置套餐,DeepSeek 峰谷调价,苹果重回 5 万亿等 shows peak/off-peak API pricing turning DeepSeek into a normal routing decision rather than an always-cheapest default.
- E249|Token经济转点:OpenClaw、Hermes到本地自研的Agent进化之路 gives a source-scoped local-model workflow for batch summaries, memory work, and repeated processing.
- Vol. 167 Token 如流水,Agent 似朝阳 and Vol. 171 假如我们有无限 Token place DeepSeek inside cost-aware routing among frontier, local, and cheaper models.
Organizational, geopolitical, and market effects
- 176: 姚顺宇,来到腾讯300天 describes DeepSeek’s breakout as an organizational shock that helped Tencent reconsider its frontier-model team.
- EP57 美股动荡,东升西降?这回是走是留 and 133.全球宏观和资本市场2025年中盘点:中国的三个温差和美国的三个预期差 connect DeepSeek to AI-capex doubt, U.S. technology valuation, and China-tech repricing.
- 把 AI 吹成核武器的人,亲手拉下了新冷战铁幕 places it in the open-model substitution response to access restrictions and export controls.
- 别在国内卷了,去美国看看只要产品好就有人付费的市场 provides a source-scoped cross-border adoption signal based on cost and utility.
Qualifications
- The sub-$6-million R1 figure is not a verified full-cost accounting and may exclude prior research, data, labor, capital expense, failed runs, post-training, and deployment.
- Benchmark parity in one source does not establish equal reliability, safety, latency, task coverage, or enterprise readiness.
- Distillation is a standard technique, but the bounded sources do not prove public accusations about unauthorized DeepSeek use of closed-model outputs.
- Several later model and product names, prices, timelines, and performance reports come from podcast hosts and may contain transcription or rumor-level uncertainty.
- Self-hosting open weights can reduce API data exposure but does not remove provenance, security, censorship, compliance, or operational risk.
What Changed
- R1’s place in the profile is now explicitly an efficiency and expectation shock, not validation of one headline cost figure.
- Enterprise self-hosting is now distinguished from sending sensitive queries to a provider-controlled API.
- The current profile more clearly separates DeepSeek’s open artifact and ecosystem effects from the full, still partly opaque model-development process.
Relationships
- Chinese Open-Weight AI Strategy - strategic context for downloadable weights, adoption, and geopolitical accessibility.
- Scaling Efficiency - engineering frame for capability per unit of compute and cost.
- Model-Infra Co-Design - relationship between DeepSeek architectures and serving infrastructure.
- Open Source AI Models - broader ecosystem category that DeepSeek helps shape.
- Model Distillation / 模型蒸馏 - technical method and contested provenance issue associated with R1-era progress.
- AI Inference Cost Structure - economics behind local use, API pricing, and routing.
- AI Export Controls - policy pressure that can constrain hardware while inducing alternative engineering routes.
- Data Sovereignty - enterprise deployment concern separating hosted API use from controlled inference.
- Nvidia - compute supplier and market symbol affected by the R1 efficiency narrative.
Sources
32 source notes across 12 shows
- Vol. 172 Codex 卖重置套餐,DeepSeek 峰谷调价,苹果重回 5 万亿等 枫言枫语
- E249|Token经济转点:OpenClaw、Hermes到本地自研的Agent进化之路 硅谷101
- Vans、匡威风光不再,经典帆布鞋为什么卖不动了? 声动早咖啡
- 179: 蒸馏风暴:一场无人公开谈论的技术竞赛 晚点聊 LateTalk
- Vol. 171 假如我们有无限 Token 枫言枫语
- 巴黎水和圣培露还能赚钱,雀巢为何要剥离水业务? 声动早咖啡
- 从蒸馏到合成数据到 RSI,模型竞争的下一个焦点是什么?|对谈 Evolvent AI 联创孟繁青 42章经
- 贾扬清:我所经历的「人工智能已死」到「AI 颠覆世界」的数年巨变丨串台「声东击西」S10E24 What's Next|科技早知道
- 177: 详解Kimi K3:强到冲击Anthropic估值的模型什么样? 晚点聊 LateTalk
- E246|何谓蒸馏?聊聊硅谷如何看中国开放模型逼近前沿 硅谷101
- 176: 姚顺宇,来到腾讯300天 晚点聊 LateTalk
- 148. 对游凯超3小时访谈:开源Infra、和模型Co-design 、“如果vLLM失败,我们会后悔一辈子” 张小珺Jùn|商业访谈录
- 174. 我们还能给算法当多久的品味老师?|对谈亚马逊AGI查晟 起朱楼宴宾客
- 172.全球宏观和资本市场2026半年度复盘与展望:AI叙事的下一步 起朱楼宴宾客
- 160.如何应对中国资产牛市的“调整期”|新书分享会成都场实录 起朱楼宴宾客
- 152.关于2026年的四个猜想 起朱楼宴宾客
- vol.124.信息过载后如何保持冷静? | 投资账复盘 起朱楼宴宾客
- Vol.114 AI的2025和DeepSeek们的未来 | 对谈复旦张奇教授 起朱楼宴宾客
- 阿里千问离职余震,在几万人的铁球里如何体面生存 科技乱炖
- 从QQ会员到豆包包月,中国人为什么总觉得软件该免费 科技乱炖
- 71. 编程的内燃机时代 内核恐慌
- EP57 美股动荡,东升西降?这回是走是留 一劳永逸
- EP58 业绩平平,也要认真"摸鱼" 一劳永逸
- 把 AI 吹成核武器的人,亲手拉下了新冷战铁幕 科技乱炖
- Vol. 167 Token 如流水,Agent 似朝阳 枫言枫语
- 当华为抛出韬定律,我们该信它到哪一步? 科技乱炖
- 别在国内卷了,去美国看看只要产品好就有人付费的市场 科技乱炖
- 138. 对罗福莉3.5小时访谈:AI范式已然巨变!OpenClaw、Agent范式很吃后训练、卡的分配、组织平权 张小珺Jùn|商业访谈录
- Vol.111 关于2025年的四个猜想 起朱楼宴宾客
- 133.全球宏观和资本市场2025年中盘点:中国的三个温差和美国的三个预期差 起朱楼宴宾客
- AI 发展了 4 年,把应用发展没了?|AI 年中复盘 42章经
- EP 34: DeepSeek R1 vs GPT-4: The $6M Model That Changed AI Economics Data Science With Sam