151. 17岁被2026年ICML收录论文的小少年:我bet开心!开心!开心!
Summary
This 张小珺Jùn|商业访谈录 episode interviews 苏廷昊, a 2009-born high-school student in Hong Kong whose attention-mechanism paper was accepted by ICML 2026. The source follows his path from ChatGPT, public courses, papers, and small-model experiments into Transformer research, then turns toward AI-native schooling, career anxiety, companion risk, AGI speculation, and happiness as a life anchor.
Key Claims
- 苏廷昊 attends 香港德瑞国际学校, was born in Shanghai, moved to Hong Kong as a child, and is presented as a concrete case of AI-native youth research rather than a conventional lab apprenticeship.
- His AI entry path ran through early OpenAI demos, ChatGPT, Andrew Ng courses on Bilibili, Andrej Karpathy’s from-scratch GPT/Transformer video, and a self-imposed 30-day paper-reading challenge.
- He originally wanted to train a small but strong language model, but GPU and token-cost estimates shifted him toward architecture experiments inside existing Transformer frameworks.
- The source says his paper work cost roughly RMB 30,000 in experiments, involved about 200-250 attempted improvements, and included a failed checkpoint setup that wasted about RMB 1,500.
- His source-scoped technical idea, attention projection residuals, starts from value residual learning, adds RMS norm on residual paths, extends the idea to key and query projections, and uses a widened first-layer projection split between current-layer use and later residual use.
- He distinguishes his work from Attention Residues in Kimi K3, treating the Kimi method as more strongly validated at large scale.
- The episode argues that a student can understand much of GPT/Transformer mechanics with matrix multiplication and high-school math, but that real intuition about training, Q/K/V, and KV cache requires long exposure and repeated experiments.
- The school-use section extends AI Default Learning Environment: classmates already use AI for homework and projects, while Su uses ChatGPT, DeepSeek, Kimi K3, and tools such as Anki to shift time away from low-value work toward self-directed learning.
- Su says AI can widen student differences: motivated students can use it to go much farther, while students who use it only to finish assignments faster may lose practice and self-control.
- The ICML social section connects Research Taste, peer community, and AI alignment: he asked dozens of attendees whether AI could cause human extinction, while also saying the race is hard to stop because states and companies have incentives to continue.
- The career section turns AI from a tool into AI Job Security Anxiety and AI Existential Meaning Anxiety: classmates worry about long professional training paths such as medicine, while Su describes weeks of low mood after reading about extinction risk and concentrated AI power.
- The companion section links Character AI, AI idol chat, teen reliance on chatbots, and a source-reported suicide-lawsuit case to AI Companion Attention Risk and Teen Chatbot Mental Health Risk.
- The closing life frame is not anti-AI. Su’s practical response is to use AI where useful, keep meaningful and happiness-producing work for himself, and make himself and people he likes happier.
Key Quotes
“开心” — the source’s recurring shorthand for Su’s life target after AI-era anxiety.
“斗不过就加入” — his pragmatic posture toward AI development and work.
“吃一堑长一智” — his mother’s response after the failed checkpoint incident.
Connections
- Su Tinghao / 苏廷昊, ICML, German Swiss International School / 香港德瑞国际学校, Andrew Ng, and Andrej Karpathy — central person, venue, school context, and learning-resource figures.
- AI-Native Youth Research, Self-Directed Learning, AI Default Learning Environment, AI Shortcut Risk, and AI Use Pacing — education and learning branch.
- Transformer Architecture, Attention Projection Residuals, Attention Residues, Kimi K3, AI Research Feedback Compression, and Research Taste — technical research branch.
- OpenAI, ChatGPT, DeepSeek, Kimi, Kimi K3, Anthropic, Gemini, and AGI Narrative — model-company and frontier-model context.
- AI Job Security Anxiety, AI Fatalistic Acceleration, AI Alignment Governance, AI Existential Meaning Anxiety, and Human Agency Under AI — risk, labor, safety, and meaning branch.
- Character AI, AI Friend Products, AI Companion Attention Risk, Teen Chatbot Mental Health Risk, and Human Connection Under AI — companion and relationship branch.
- AI For Fun, Youth Happiness After Growth, Meaning Through Experience, and Embodied Intelligence / 具身智能 — happiness, embodied life, and future-society branch.
Contradictions
- No direct contradiction found.
- The source qualifies technical claims by noting that some names and details come from automatic transcription; dataset name, data scale wording, and some model-release references should remain source-scoped until checked against the actual paper or primary materials.
- The source reinforces existing AI-anxiety pages, but keeps extinction, unemployment, UBI, and AGI-timeline comments as one student’s views rather than settled forecasts.