2026-09-24 日报
主题: 编码一次多处复用:推荐级联共享主干与 KV 不变扩参
标签: parameter-scaling · inference-serving · moe · multi-task · industrial
📊 统计: 共 6 篇 · 精读 2 · 🏢 工业界 2 · 🎓 学术 4 · generative-rec 1 · discriminative-rec 3 · llm 2 · other 1
综述
今日 6 篇(判别式推荐 3、LLM 2、生成式推荐 1、其他 1),工业署名 2 篇,精读 2 篇。字节 OneTrans-V2 把召回、粗排、精排作为一个 Transformer 的三个任务联合训练,但服务仍跑三级级联,共享的只是一次序列编码、一套参数和一个训练作业,是共享主干而非替换级联。线上 GMV/u +9.74% 的对照是整条旧级联,混合了 2.4× 算力、MoE/μP 等 scaling 配方、DCGR、联合训练与蒸馏,无分解;3.2× QPS 中共享仅占 1.53×,联合训练约占离线边际 1/7。StepFun 的 KITE 中途把 33.8B MoE 扩成 67B 两塔,Decoder 只读 Prefiller KV;等 FLOPs 下游全领先,但 NLL 与 63B 持平,推理节省仅为估算、无消融。两篇都靠“编码一次、多处复用”摊薄成本,但收益多来自 scaling 与生长,统一和 KV 不变本身的贡献仍待隔离。
重点论文
OneTrans-V2 · ⭐ 7/10
🏢 ByteDance · 生成式推荐 / 判别式推荐
OneTrans-V2 把召回、粗排、精排作为一个因果 Transformer 的三个任务联合训练,共享一次编码的用户序列上下文并保留各阶段原生候选特征,辅以决策前缀+业务偏置的 DCGR 召回、fine→pre 模型内蒸馏、MoE/μP scaling 与 SNT 训练组织,线上替换三模型级联 GMV/u +9.74%、同硬件 3.2x QPS。
KITE · ⭐ 6/10
KITE: KV-Invariant Transformer Expansion for Efficient Agentic LLM Scaling
🏢 StepFun · LLM
StepFun 的 KITE 在训练中途把小模型扩成两塔:源模型保留为唯一产 KV 的 Prefiller,复制出只读逐层 KV 的 Decoder 联合续训,使 bulk prefill 保持源模型尺寸;SST 实例(33.8B→67B MoE)在等理论训练 FLOPs 下训练 loss 1.5900 优于 47B/63B 从头训练(1.6006/1.5921)、7 项下游全领先而 held-out NLL 与 63B 持平,75:25 P:D 下参数量代理推理代价低 6.7%/31.6%,但无实测 serving、无生长/KV 复用基线、无消融。
RefineICL · ⭐ 5/10
What Do Tabular Foundation Models Compute In Context? In-Situ Representation Refinement through Attention-Gated Updates
🎓 学术 · 其他
将表格基础模型的上下文学习解释为“原位表示精炼”:支持集标签引导 episode 内表示更新并迁移到查询;据此设计注意力门控、无 FFN 的 RefineICL,在 TabArena 上比 TabPFN-3 高 31.4 Elo,且去掉 FFN 不损精度、节省推理内存。
全部论文
| 模型 | 标题 | 类别 | 公司 | 摘要分 | 精读分 |
|---|---|---|---|---|---|
| OneTrans-V2 | OneTrans-V2: Unifying Retrieval, Pre-rank, and Fine-rank with One Transformer in Industrial Recommender | 生成式 / 判别式 | 🏢 ByteDance | 8 | 7 |
| KITE | KITE: KV-Invariant Transformer Expansion for Efficient Agentic LLM Scaling | LLM | 🏢 StepFun | 7 | 6 |
| RefineICL | What Do Tabular Foundation Models Compute In Context? In-Situ Representation Refinement through Attention-Gated Updates | 其他 | 🎓 学术 | 5 | — |
| — | The Capability Manifold and ML Scaling Laws | LLM | 🎓 学术 | 4 | — |
| — | A Flexible Recommendation System for Individuals and Groups | 判别式 | 🎓 学术 | 4 | — |
| — | A Systematic Benchmark of Explainable Methods for Temporal Attribution in Sequential Recommendation Systems | 判别式 | 🎓 学术 | 4 | — |