Writing
Research notes, technical essays, and system-level thinking
Original essays around efficient LLMs, long-context systems, agent engineering, model compression, and AI-native product interfaces.
Latest Essay
On-Policy Distillation
Training Students on Their Own Mistakes
A practical guide to on-policy distillation: why student-generated trajectories reduce exposure bias, how the main OPD families differ, and where the method fits alongside SFT a...
On-Policy Distillation
Training Students on Their Own Mistakes
A practical guide to on-policy distillation: why student-generated trajectories reduce exposure bias, how the main OPD families d...
Trace Data Foundry:当 Agent 执行轨迹成为可交易的模型能力资产
把 Agent 执行轨迹视为能力资产:分析高价值 trace 的供给、验证、定价与交易机制,并推演 Trace Data Foundry 的产品边界。
你和 AI 的关系,决定了你的能力走向
区分 Skill 化、Harness 化与超长周期 Agent:专家能力如何被封装、约束和持续学习,以及人在其中如何避免能力被动退化。
DeepSeek-V4 论文解读:百万 Token 上下文不是窗口竞赛,而是系统工程
DeepSeek-V4 Paper Notes: Million-Token Context as Systems Engineering
从系统工程视角解读 DeepSeek-V4:百万 Token 上下文如何同时依赖稀疏注意力、KV Cache 压缩、通信隐藏与后训练协同。
长上下文的终局,不是更大的窗口
长上下文竞争正在从窗口大小转向状态连续服务:真正关键的是服务寿命、外部记忆、状态恢复、工具验证与长期推理成本。
From Ultra-Long Context to Lifelong Service: Why Open Models Are Likely to Become State-Centric Serving Systems
Why the long-context race is becoming a state-serving problem involving service lifetime, external memory, recovery, verification...
Claude Code 与 Agent Harness 阅读地图
A Reading Map for Claude Code and Agent Harnesses
一张理解 Claude Code 与 Agent Harness 的分层阅读地图:从智能体循环和 App Server 出发,再进入上下文、工具、长任务与多智能体系统。
Harness Engineering:把会写代码的模型,变成真正能交付的软件系统
Harness Engineering 关注的不是模型能否写出第一版代码,而是如何用环境、验证、权限、记忆和反馈回路把概率性输出变成持续交付。
The Death of the App: Why the "Intent Canvas" is the Endgame of Operating Systems
A systems blueprint for moving beyond isolated apps toward an intent-driven canvas assembled by asynchronous agents and reliable ...
App 的尽头是"画布":零应用时代的底层架构逻辑
从 App 孤岛走向意图驱动画布:分析生成式界面为什么不能只靠模型现场写代码,以及异步多智能体 UI 引擎需要哪些基础层。
大模型对话格式全景
把 Chat Template 视为模型输入协议,解释角色、工具调用、特殊 token 和训练阶段边界如何共同决定对话系统的正确性。
LLM Agent 记忆管理方案
系统梳理 LLM Agent 的记忆管理方案,从工作记忆、长期记忆和检索机制,到写入、更新、遗忘与评测的完整生命周期。
GPT-OSS Model Card 解析
围绕 GPT-OSS 的模型规格、架构、推理能力、安全评估和部署方式展开,帮助读者从 Model Card 中提取真正影响使用决策的信息。
原生轻量化大语言模型
Native Small Language Models
从架构、数据、训练和部署四个层面讨论原生小语言模型,说明轻量化不只是压缩大模型,而是一套面向效率的协同设计。
可验证奖励的强化学习(RLVR)
Reinforcement Learning with Verifiable Rewards (RLVR)
解释可验证奖励强化学习如何利用数学、代码等可自动检查的反馈训练推理模型,并梳理 GRPO、奖励设计、训练流程与常见风险。
大型语言模型量化技术:原理、前沿与实践
LLM Quantization: Principles, Frontiers and Practice
从量化误差与数值表示出发,系统梳理 PTQ、QAT、混合精度和 QLoRA,并给出大语言模型压缩与部署的实践路径。
No matching essays.