90 papers from 2025–2026 testing hard Semantic IDs against textual addresses, summaries, generated retrieval vectors, soft tokens, hybrid systems, and ID-free representations · 90 papers · 中文 ↗
Direct answer. Yes, recent systems replace Semantic-ID inputs and outputs with controlled attribute strings, editable keyword paths, structured descriptions, task-specific summaries, flow-generated target embeddings, autoregressively generated retrieval vectors, and soft feature tokens. The strongest criticism is also broader now: SID systems can saturate under scaling, lose reachable items through constrained decoding, fail to reach future cold items, amplify popularity, and depend on brittle decoding priors. The useful question is not whether SIDs are universally good or bad, but which interface still wins after item-level grounding, temporal reachability, perturbation, maintenance, and strong text/vector baselines are matched.
0 / 90 read
Decision rule. Use a hard SID only when compact constrained decoding over a known catalog is itself required, one-item grounding or bucket resolution is explicit, and the address can be maintained under catalog change. Prefer controlled text when the model must read, edit, or explain meaning; continuous vectors when insertion and ANN retrieval matter; soft or hybrid tokens when heterogeneous signals must remain differentiable; and ID-free content representations when cold start matters more than memorized identity. If a dense scorer, textual path, or generated vector recovers the gain without collision and refresh machinery, the simpler non-SID interface is the better default.