When Copying Is Hard: Copy-Constrained Decoding for Exact Span Reproduction
Efficiency
Trustworthiness
Applications
NeurIPS 2026 CopyGen is a lightweight decoding-time plugin that makes LLMs reproduce source spans more exactly and faster, improving average exact match by 35% relative without finetuning. From Individuals to Crowds: Dual-Level Public Response Prediction in Social Media
Personalization
Applications
ACM MM 2025 Proposed a personalized response prediction framework with PAC-LoRA, enabling user-specific comment generation while modeling crowd-level sentiment trends. LITERARYBIGFIVE: Author-Personalized Text Generation in a Unified Interpretable Space
Personalization
Interpretability
EMNLP 2026 Findings We reframe different writing patterns into a unified interpretable style space, enabling personalized author-aware text generation framework via steering outputs toward target authorial characteristics. The Cylindrical Representation Hypothesis for Language Model Steering
Interpretability
ICML 2026 [Invited Talk @ Cohere] Introduced CRH, a cylindrical geometric model to explain instability in activation steering, and empirically verified its sample-specific structure across models, concepts, and steering methods. M3MAD-Bench: Are Multi-Agent Debates Really Effective Across Domains and Modalities?
Applications
ACM MM 2026 M3MAD-Bench is a standardized multimodal benchmark for assessing multi-agent debate methods across domains, modalities, and performance-cost trade-offs. TransPrune: Token Transition Pruning for Efficient Large Vision-Language Model
Efficiency
CVPR 2026 TransPrune is a training-free token pruning method that combines Token Transition Variation with instruction-guided attention to reduce LVLM inference cost by over 50% with comparable performance.
|