← 返回报告列表

2026-06-26 日报

日报 📅 2026-06-25
Agent 驱动工业推荐自迭代与 model-centric 扩参双线
industrial pretrained-lm parameter-scaling transformer
📊 共 6 篇 · 精读 4

2026-06-26 日报

主题: Agent 驱动工业推荐自迭代与 model-centric 扩参双线

标签: industrial · pretrained-lm · parameter-scaling · transformer

📊 统计: 共 6 篇 · 精读 4 · 🏢 工业界 3 · 🎓 学术 3 · llm 2 · discriminative-rec 2 · generative-rec 1 · other 1

综述

今日共 6 篇相关论文,4 篇精读、2 篇略读;类别上 LLM 与判别式推荐各 2 篇、生成式推荐 1 篇、其他 1 篇,工业界(快手×2、腾讯)与学术界各占其半。UniFormer(快手)把工业推荐建模空间显式拆为特征空间 FIM 与任务空间 TIM 分别堆叠扩参,将 scaling 范式从组件级推进到 model-centric。NOVA(腾讯)提出“验证感知”LLM 多智能体框架,把架构演进形式化为受 SGD 启发的“架构梯度”搜索加多阶段验证级联,在训练前拦截“能跑但结构无效”的静默失败,线上 GMV +1.25%~2.02%。AgentX(快手)则把 idea-to-launch 闭环整体自动化(Brainstorm→Developing→Evaluation→SGPO 自进化),三周将 374 个想法转为 10 个上线。MO-DiT+HPPO(北大)用球面 flow-matching 扩散 transformer 做 pattern-preserving 属性检索,并以 HPPO 对齐在线 Joint@K。整体看,LLM agent 正从辅助工具走向工业推荐系统自动演进的主循环,scaling 重心转向 model-centric,生成式与扩散建模持续向召回侧渗透。

重点论文

UniFormer · ⭐ 8/10

UniFormer: Efficient and Unified Model-Centric Scaling for Industrial Recommendation

🏢 Kuaishou · 判别式推荐

提出 UniFormer,把工业推荐的建模空间显式拆为特征空间(FIM)与任务空间(TIM)分别堆叠扩参,配语义化 tokenization 做用户-物品解耦加速、多序列 cross-attention 防偏好坍塌、多视角 FFN 灵活分配容量,将 scaling 范式从组件级/特征空间协同推进到 model-centric。

NOVA · ⭐ 8/10

NOVA: A Verification-Aware Agent Harness for Architecture Evolution in Industrial Recommender Systems

🏢 Tencent · LLM

NOVA 是腾讯提出的'验证感知'LLM 多智能体框架,把工业推荐架构演进形式化为受 SGD 启发的'架构梯度'搜索 + 多阶段验证级联(训练前拦截'能跑但结构无效'的静默失败并转为禁止方向)+ L1-L4 分级与 AutoRun/Copilot 管控,在腾讯广告 L3 文献到生产任务上 EPR 达 60%、人工耗时缩短 13.5×,线上 A/B GMV +1.25%~+2.02% 且 pCVR bias 同步下降。

AgentX · ⭐ 8/10

AgentX: Towards Agent-Driven Self-Iteration of Industrial Recommender Systems

🏢 Kuaishou · LLM

快手 AgentX 是生产部署的多 agent 系统,把工业推荐的 idea-to-launch 闭环(Brainstorm 提案 → Developing 双轨编码 → Evaluation guardrail-veto A/B → SGPO 语义梯度自进化)整体自动化,三周三 worker 把 374 想法转 10 个上线、主 feed +0.561% app 时长、生活服务年化收入超 1 亿元。

MO-DiT+HPPO · ⭐ 7/10

Generative Retrieval via Diffusion Transformer with Metric-Ordered Sequence Training and Hybrid-Policy Preference Optimization

🎓 学术 · 生成式推荐

提出 MO-DiT+HPPO 解决 pattern-preserving attribute retrieval:用球面 flow-matching 的扩散 transformer 读 item embedding 序列生成连续 query 做最近邻检索,以 metric-ordered 序列把稀疏在线检索标签变成同模式低→高密度轨迹训练 CPT+tail-centroid SFT,再用 HPPO(在线 Joint@K 作 reward 的 DPO 式偏好优化+混合候选池+反 reward-hacking 的 Pareto pair filter)对齐真实在线交集指标,在四个内部域上把 Joint@K 显著推高。

全部论文

模型 标题 类别 公司 摘要分 精读分
UniFormer UniFormer: Efficient and Unified Model-Centric Scaling for Industrial Recommendation 判别式 🏢 Kuaishou 8 8
NOVA NOVA: A Verification-Aware Agent Harness for Architecture Evolution in Industrial Recommender Systems LLM 🏢 Tencent 7 8
AgentX AgentX: Towards Agent-Driven Self-Iteration of Industrial Recommender Systems LLM 🏢 Kuaishou 7 8
MO-DiT+HPPO Generative Retrieval via Diffusion Transformer with Metric-Ordered Sequence Training and Hybrid-Policy Preference Optimization 生成式 🎓 学术 7 7
TRUST TRUST: Item-Calibrated Interval Evidence for Temporal Session-Based Recommendation 判别式 🎓 学术 4 —
— Sketched Linear Contrastive Learning: Approximation, Optimization, and Statistical Scaling 其他 🎓 学术 4 —