快一点!再快一点!快到世界能实时生成|和生数科技张金涛聊:Vidu S1、推理加速、实时交互视频
Summary
This Shizilukou Crossing episode interviews 张金涛 / Zhang Jintao of 生数科技 / Shengshu Technology about Vidu S1, Real-Time Interactive Video Generation, and the engineering work needed to make video generation faster than playback. The source argues that real-time visual generation is not just a bigger-hardware problem: SAGE Attention, TurboDiffusion, model distillation, sparse attention, cluster scheduling, and hardware-algorithm co-design all sit inside an Inference Acceleration Stack. Its product synthesis is that streaming, interactive video moves Video Models from offline clip generation toward live sessions where characters must understand voice, video, screens, and user intent while preserving speed, identity, and coherence.
Key Claims
- 张金涛 / Zhang Jintao is described as a 26-year-old [[TsinghuaUniversity|Tsinghua University]] doctoral student who had previously done inference-acceleration research and visited [[UCBerkeley|UC Berkeley]].
- At 生数科技 / Shengshu Technology, Zhang is said to work on Vidu S1, inference acceleration, cluster deployment, and the full Streaming Video Generation chain from data and training algorithms to engineering and evaluation.
- Vidu S1 is presented as a live interactive product: a user can upload a character image, choose or clone a voice, and talk to the character through voice while the generated video responds in real time.
- The source-scoped performance claim is that S1 can stream long-form generation at 540P and roughly 25 to 42 FPS, with the team prioritizing latency and responsiveness before maximum image quality.
- SAGE Attention is described as a faster Attention operator that matters more for multimodal and video generation because those workloads can be compute-bound rather than only memory-bound.
- TurboDiffusion is presented as a model-level acceleration project that combines faster operators with distillation and sparse-attention fine-tuning to reduce diffusion-model complexity.
- Zhang divides inference acceleration into operator acceleration, model-complexity reduction, and engineering/deployment optimization, including multi-card parallelism, communication-compute overlap, request scheduling, and large-cluster deployment.
- The source distinguishes Streaming Video Generation from ordinary offline video generation: every frame must be produced in time, quality must not drift over long sessions, and input feedback must remain correct.
- Zhang frames online visual entertainment as a larger long-term demand category than pre-generated clips, with possible scenes including conversation, romance, games, pets, desktop assistance, and everyday companion experiences.
- The source attributes part of Chinese video-model strength to visual-entertainment data quantity, data quality, preference alignment, data construction, and ecosystems such as short video, livestream commerce, and social platforms.
- Zhang argues that AI is not merely a bubble because it can meet many human needs, but he also says technical strength alone does not make someone a good founder; management, coordination, and judgment matter.
Key Quotes
“生成速度必须超过播放速度” — Zhang’s concise boundary for streaming video generation.
“在线的、实时交互的视频” — the source’s main product category distinction.
“每一行代码都正确” — Zhang’s engineering definition of world-leading execution.
Connections
- 张金涛 / Zhang Jintao, 生数科技 / Shengshu Technology, Vidu, and Vidu S1 — guest, company, model/product family, and real-time product case.
- SAGE Attention, TurboDiffusion, and Inference Acceleration Stack — technical acceleration thread.
- Streaming Video Generation, Real-Time Interactive Video Generation, Video Models, and Multimodal Intelligence — model and interaction frame.
- AI Interactive Entertainment, AI Simulation Content, AI Startup Unit Economics, and Product Led Willingness To Pay — user-demand and monetization frame for live visual sessions.
- AI Inference Cost Structure, MaaS Infrastructure, High-Throughput Inference Batching, Low-Latency Inference Chip, and AI Chip Specialization — serving, scheduling, and hardware-software economics.
- GPU, Nvidia, AMD, Huawei, ByteDance, Tencent, and Google — hardware and company context mentioned in the source’s adoption and deployment discussion.
- Tsinghua University / 清华大学, UC Berkeley, Transformer Architecture, Mixture of Experts, and Diffusion Transformers — research and architecture context.
Contradictions
- No direct contradiction found. The source extends the existing Video Models and World Models pages by separating offline video quality from real-time interactive continuity, while keeping S1’s performance, price, and adoption claims source-scoped rather than independently verified.