Research Plan

Reasoning over Semantic IDs

71 papers from 2025–2026 separating explicit CoT, latent computation, adaptive routing, process credit, grounded retrieval, and verifier-guided SID decoding · 71 papers · 中文 ↗
Scope rule. Reasoning counts only when intermediate computation, evidence, reward, or search changes the item-token decision. Ordinary DPO, listwise alignment, model averaging, and test-time parameter adaptation have been removed. Some non-SID latent methods remain as boundary comparisons because they isolate a reasoning mechanism that SID systems can test. The central question is no longer simply whether to think, but where to attach credit, when extra computation is needed, and how to stop it from overriding behavioral evidence.
0 / 71 read
Decision rule. Start from the strongest non-thinking generator. Add process credit when SID prefixes fail, latent steps when fixed encoders miss interactions, adaptive routing when difficulty varies, external tools when evidence is missing, and a verifier when beam search prunes valuable branches. Keep the reasoning module only if the gain survives equal-compute baselines and justifies latency, extra tokens, and new failure modes.