Research Plan

LLM and Synthetic Data Augmentation for Recommendation

57 papers from 2025–2026 on cleaning interaction logs, grounded and cross-domain augmentation, structured synthetic curricula, and user simulation · 57 papers · 中文 ↗
Scope: 2025–2026, LLM-centered with data-centric boundary papers. The page focuses on methods that modify, synthesize, or transfer recommendation training data. Most use LLMs as offline judges, annotators, semantic generators, or simulators; a small number of non-LLM methods such as SCALR are retained when they directly adapt the synthetic-data paradigm or provide a necessary data-generation comparison. Treat every generated edge or deleted event as a hypothesis rather than truth.
0 / 57 read
Deployment rule. Preserve the raw log, attach calibrated soft weights or confidence labels, and distill expensive LLM decisions into a reusable lightweight cleaner whenever possible. Report recommendation accuracy together with long-tail quality, calibration, data efficiency, and results on untouched real interactions.