On the Effectiveness-Fluency Trade-Off in LLM Conditioning: A Systematic Study
On the Effectiveness-Fluency Trade-Off in LLM Conditioning: A Systematic Study - Apple Machine Learning Research research area Methods and Algorithms , research area Speech and Natural Language Processing conference EMNLP content type paper published September 2026 On the Effectiveness-Fluency Trade-Off in LLM Conditioning: A Systematic Study Authors Iuri Macocco†, Pau Rodríguez Lopez, Arno Blaas, Luca Zappella, Marco Baroni†*, Xavier Suau...

On the Effectiveness-Fluency Trade-Off in LLM Conditioning: A Systematic Study - Apple Machine Learning Research research area Methods and Algorithms , research area Speech and Natural Language Processing conference EMNLP
content type paper published September 2026
On the Effectiveness-Fluency Trade-Off in LLM Conditioning: A Systematic Study
Authors Iuri Macocco†, Pau Rodríguez Lopez, Arno Blaas, Luca Zappella, Marco Baroni†, Xavier Suau Cuadros
Controlling the output of Large Language Models (LLMs) is a central challenge for their reliable deployment, yet a clear understanding of the involved trade-offs remains elusive. Current approaches to conditioning are often evaluated with a narrow focus on their effectiveness at injecting or removing a target concept, neglecting generation quality. We systematically investigate a range of conditioning methods in both injection and removal scenarios. We find that efficient steering methods frequently achieve conditioning at a steep cost to fluency. Furthermore, we identify a critical yet previously overlooked interaction with the training paradigm: activation steering methods are far less effective on instruction-tuned models than on their base counterparts. Simple prompting and full-fledged supervised fine-tuning, on the other hand, are viable options for concept injection, but are not as good at concept removal. Finally, cheaply computed textual metrics highly correlate to costly LLM-as-judge scores, and provide insights on the behavior of conditioning methods.
September 18, 2026 research area Human-Computer Interaction , research area Methods and Algorithms Transactions on Machine Learning Research (TMLR)
Activation steering has emerged as a powerful method for guiding the behavior of generative models towards desired outcomes such as toxicity mitigation. However, most existing methods apply interventions uniformly across all inputs, degrading model performance when steering is unnecessary. We introduce Dynamically Scaled Activation Steering (DSAS), a method-agnostic steering framework that decouples when to steer from how to steer. DSAS…
STEER: Semantic Turn Extension-Expansion Recognition for Voice Assistants
November 8, 2023 research area Speech and Natural Language Processing conference EMNLP
In the context of a voice assistant system, steering refers to the phenomenon in which a user issues a follow-up command attempting to direct or clarify a previous turn. We propose STEER, a steering detection model that predicts whether a follow-up turn is a user’s attempt to steer the previous command. Constructing a training dataset for steering use cases poses challenges due to the cold-start problem. To overcome this, we…
Discover opportunities in Machine Learning.
Our research in machine learning breaks new ground every day.
Sources
Related stories

Model Auditing: Faithfulness of Post-hoc Explanations Investigated
A study published on arXiv cs.CV investigates the faithfulness of post-hoc explanations for a pedestrian detection model across different domains.

DeskForge Corpus Improves Computer-Use Agent Accuracy
A new corpus of annotated desktop observations, DeskForge, enables computer-use agents to achieve higher accuracy in complex desktop scenes.

edithero benchmark for long-horizon 3d editing
The EditHero benchmark is introduced to evaluate and improve the reliability of long-horizon, part-level 3D editing.

ComfyUI v0.39.0 Release Brings New Features and Improvements
The ComfyUI v0.39.0 release includes dynamic group widgets, new model nodes, and resolved minimax vae offload issue.