2026-06-26 日报
主题: Agent 驱动工业推荐自迭代与 model-centric 扩参双线
标签: industrial · pretrained-lm · parameter-scaling · transformer
📊 统计: 共 6 篇 · 精读 4 · 🏢 工业界 3 · 🎓 学术 3 · llm 2 · discriminative-rec 2 · generative-rec 1 · other 1
综述
今日共 6 篇相关论文,4 篇精读、2 篇略读;类别上 LLM 与判别式推荐各 2 篇、生成式推荐 1 篇、其他 1 篇,工业界(快手×2、腾讯)与学术界各占其半。UniFormer(快手)把工业推荐建模空间显式拆为特征空间 FIM 与任务空间 TIM 分别堆叠扩参,将 scaling 范式从组件级推进到 model-centric。NOVA(腾讯)提出“验证感知”LLM 多智能体框架,把架构演进形式化为受 SGD 启发的“架构梯度”搜索加多阶段验证级联,在训练前拦截“能跑但结构无效”的静默失败,线上 GMV +1.25%~2.02%。AgentX(快手)则把 idea-to-launch 闭环整体自动化(Brainstorm→Developing→Evaluation→SGPO 自进化),三周将 374 个想法转为 10 个上线。MO-DiT+HPPO(北大)用球面 flow-matching 扩散 transformer 做 pattern-preserving 属性检索,并以 HPPO 对齐在线 Joint@K。整体看,LLM agent 正从辅助工具走向工业推荐系统自动演进的主循环,scaling 重心转向 model-centric,生成式与扩散建模持续向召回侧渗透。
重点论文
UniFormer · ⭐ 8/10
UniFormer: Efficient and Unified Model-Centric Scaling for Industrial Recommendation
🏢 Kuaishou · 判别式推荐
提出 UniFormer,把工业推荐的建模空间显式拆为特征空间(FIM)与任务空间(TIM)分别堆叠扩参,配语义化 tokenization 做用户-物品解耦加速、多序列 cross-attention 防偏好坍塌、多视角 FFN 灵活分配容量,将 scaling 范式从组件级/特征空间协同推进到 model-centric。
NOVA · ⭐ 8/10
🏢 Tencent · LLM
NOVA 是腾讯提出的'验证感知'LLM 多智能体框架,把工业推荐架构演进形式化为受 SGD 启发的'架构梯度'搜索 + 多阶段验证级联(训练前拦截'能跑但结构无效'的静默失败并转为禁止方向)+ L1-L4 分级与 AutoRun/Copilot 管控,在腾讯广告 L3 文献到生产任务上 EPR 达 60%、人工耗时缩短 13.5×,线上 A/B GMV +1.25%~+2.02% 且 pCVR bias 同步下降。
AgentX · ⭐ 8/10
AgentX: Towards Agent-Driven Self-Iteration of Industrial Recommender Systems
🏢 Kuaishou · LLM
快手 AgentX 是生产部署的多 agent 系统,把工业推荐的 idea-to-launch 闭环(Brainstorm 提案 → Developing 双轨编码 → Evaluation guardrail-veto A/B → SGPO 语义梯度自进化)整体自动化,三周三 worker 把 374 想法转 10 个上线、主 feed +0.561% app 时长、生活服务年化收入超 1 亿元。
MO-DiT+HPPO · ⭐ 7/10
🎓 学术 · 生成式推荐
提出 MO-DiT+HPPO 解决 pattern-preserving attribute retrieval:用球面 flow-matching 的扩散 transformer 读 item embedding 序列生成连续 query 做最近邻检索,以 metric-ordered 序列把稀疏在线检索标签变成同模式低→高密度轨迹训练 CPT+tail-centroid SFT,再用 HPPO(在线 Joint@K 作 reward 的 DPO 式偏好优化+混合候选池+反 reward-hacking 的 Pareto pair filter)对齐真实在线交集指标,在四个内部域上把 Joint@K 显著推高。
全部论文
| 模型 | 标题 | 类别 | 公司 | 摘要分 | 精读分 |
|---|---|---|---|---|---|
| UniFormer | UniFormer: Efficient and Unified Model-Centric Scaling for Industrial Recommendation | 判别式 | 🏢 Kuaishou | 8 | 8 |
| NOVA | NOVA: A Verification-Aware Agent Harness for Architecture Evolution in Industrial Recommender Systems | LLM | 🏢 Tencent | 7 | 8 |
| AgentX | AgentX: Towards Agent-Driven Self-Iteration of Industrial Recommender Systems | LLM | 🏢 Kuaishou | 7 | 8 |
| MO-DiT+HPPO | Generative Retrieval via Diffusion Transformer with Metric-Ordered Sequence Training and Hybrid-Policy Preference Optimization | 生成式 | 🎓 学术 | 7 | 7 |
| TRUST | TRUST: Item-Calibrated Interval Evidence for Temporal Session-Based Recommendation | 判别式 | 🎓 学术 | 4 | — |
| — | Sketched Linear Contrastive Learning: Approximation, Optimization, and Statistical Scaling | 其他 | 🎓 学术 | 4 | — |