Fine-tune a search agent with multi-turn RL on Amazon SageMaker AI
Search agents powered by large language models (LLMs) are transforming how enterprises retrieve information.

- Search agents powered by large language models (LLMs) are transforming how enterprises retrieve information.
- Environment and setup Our setup relied on the following: An AWS account with access to Amazon SageMaker AI in the US West (Oregon) AWS Region (us-west-2).
- Solution overview In this post, we use Amazon SageMaker AI MTRL to fine-tune a Qwen3.6-27B model (supported in the US West (Oregon) Region (us-west-2)) for a search agent.
Search agents powered by large language models (LLMs) are transforming how enterprises retrieve information. Rather than requiring users to craft the perfect query, a search agent autonomously decides what to search for, which retrieval strategy to use, and when to stop searching. It does this across multiple rounds of interaction, refining its approach based on what it has already retrieved. However, getting this multi-step behavior to work well is hard. No base model arrives knowing your tools or your environment. Prompt a small model and you rarely get dependable multi-turn behavior. Prompt a frontier model and it often works, but you pay for that capability in latency and cost. Fine-tuning offers a third path: you teach a small model your tools and environment directly. The result is a small model’s speed and cost with the reliability that would otherwise require a frontier model. Even though fine-tuning is the natural next step, the traditional approaches each fall short. Supervised fine-tuning (SFT) depends on expert demonstrations of ideal multi-turn trajectories, which are costly to collect and usually don’t exist for your setup. Single-turn reinforcement learning (RL), such as RL with verifiable rewards (RLVR), scores one response at a time. But a search agent makes interdependent decisions across many turns, each one building on the context before it. Optimizing a single step in isolation misses those dependencies entirely. What you need is a training approach that optimizes the agent across the full multi-turn trajectory. The reward signal only needs to reflect whether the final outcome was good. That is exactly what multi-turn reinforcement learning (MTRL) provides. It trains the agent to make good decisions across a full sequence of steps, bakes in your environment-specific behavior, and gives you control over output quality. Because you run a smaller, specialized model, you also get faster, cheaper inference. In this post, we describe how we fine-tuned a search agent using Amazon SageMaker AI multi-turn reinforcement learning (MTRL) and share the results we observed in retrieval quality and reliability. We begin by explaining what Amazon SageMaker AI MTRL is and how it works. What is Amazon SageMaker AI MTRL? With Amazon SageMaker AI MTRL, you can fine-tune LLMs using reinforcement learning in multi-turn interaction settings. It frames an agentic task as a sequence of decisions, uses multi-turn rollouts to generate training data, and optimizes the model with policy gradient algorithms. Amazon SageMaker AI MTRL offers: Modular agent-environment interface: You keep integration low-code. You define custom rewards, custom tool loops, and multi-turn conversation shapes. Serverless execution: You get production-scale agentic RL at per-token pricing without provisioning or managing GPU clusters. Asynchronous rollout and trajectory collection: You run generation and gradient updates in parallel with bounded off-policy staleness, so training stays fast without drifting too far from the current policy. A native algorithm library: You choose from Proximal Policy Optimization (PPO), Clipped Importance Sampling Policy Optimization (CISPO), and importance-sampling (IS) losses, paired with group-based advantage estimators (GRPO, GRPO pass@k, RLOO, and more). Resumable training: You can split long training runs across multiple jobs to work beyond single-job time limits. Trajectory and reward observability: You can inspect what your agent did turn by turn and across training steps in MLflow managed by Amazon SageMaker AI. Evaluation jobs: You can report reward, pass@k, and trajectory metrics before deploying to an Amazon SageMaker AI endpoint or Amazon Bedrock. This makes MTRL a natural fit for search agents: you have a clear reward signal (retrieval quality), a multi-turn interaction loop (the agent issuing queries and receiving results), and a well-defined environment (the search tools). Environment and setup Our setup relied on the following: An AWS account with access to Amazon SageMaker AI in the US West (Oregon) AWS Region (us-west-2). Training and validation datasets uploaded to Amazon Simple Storage Service (Amazon S3) in the required format (see the MTRL documentation for details). A deployed agent endpoint that exposes BM25 and vector search tools for the MTRL environment to call during rollouts. Familiarity with Amazon SageMaker AI and Python. Solution overview In this post, we use Amazon SageMaker AI MTRL to fine-tune a Qwen3.6-27B model (supported in the US West (Oregon) Region (us-west-2)) for a search agent. The search agent is an LLM-powered system that autonomously uses search tools to find, gather, and synthesize information to answer a question or complete a task. We focus on an enterprise search setting where the agent has two tools available: Lexical search (BM25): Finds exact keyword matches by counting word frequencies. Best for queries with specific terms or identifiers. Vector search: Converts queries and documents into embedding vectors and computes similarities. This is recommended for semantic or conceptual queries. We also limit the number of turns (one turn is one round of user-assistant interaction), preventing the agent from generating excessively long responses and encouraging efficient search behavior. Training setup This section walks through the three key components of our training setup: the datasets, the reward function, and the MTRL job configuration. Datasets We use the following datasets for training and testing: Dataset Train/Test Public Link Description FRAMES Train HuggingFace Multi-hop factoid QA requiring synthesis across multiple Wikipedia articles. BRIGHT Train HuggingFace Reasoning-intensive retrieval across 12 domains, where relevance needs reasoning, not keyword overlap. Enterprise RAG Train GitHub A benchmark of over 500,000 synthetic enterprise documents and 500 questions for training and evaluating RAG systems on internal company knowledge ESCI Train GitHub A multilingual product-search dataset labeling query–product pairs as Exact, Substitute, Complement, or Irrelevant Musique Train GitHub A question-answering dataset built by composing single-hop questions into challenging multi-hop reasoning problems. MLQA Train HuggingFace A parallel extractive question-answering benchmark covering seven languages for evaluating cross-lingual comprehension. FreshStack Test HuggingFace Recent developer-tech Q&A over Stack Overflow + docs for five topics. WixQA Test HuggingFace Help-center support QA over the Wix knowledge base. BrowseComp-Plus Test GitHub Hard deep-research queries over ~100K human-verified open-web documents. Wands Test HuggingFace A Wayfair product-search dataset containing over 42,000 query–product relevance judgments labeled as Exact, Partial, or Irrelevant. We preprocess the datasets into the format recognized by the MTRL service according to the documentation. We further reserve 5 percent of the training instances within each dataset as validation instances. Reward function Our main metric is nDCG@10 (Normalized Discounted Cumulative Gain at rank 10). nDCG@10 is a standard information retrieval metric that measures how well the top 10 retrieved documents match the ideal ranking. It rewards systems that place highly relevant documents near the top of the list, with a score of 1.0 meaning perfect ranking and 0.0 meaning no relevant documents retrieved. We use nDCG@10 directly as the reward function in MTRL. This is a trajectory-level reward where the agent completes its full multi-turn search, and the reward reflects how well the final retrieved documents match the ground truth. When the agent reaches the maximum number of turns or the maximum sampling tokens in a single turn, we assign a reward of -1 to explicitly teach the model to avoid these failure modes. This kind of penalty-based reward design is effective at steering the model away from undesirable behavior without needing to hand-craft complex intermediate rewards. MTRL job configuration One of the appeals of Amazon SageMaker AI MTRL is how little you need to configure to get going. We change three hyperparameters and leave everything else at the default values. The following code snippet shows how to configure and launch the training job using the MultiTurnRLTrainer SDK: from sagemaker.modules.train import MultiTurnRLTrainer trainer = MultiTurnRLTrainer( model_id="qwen3.6-27b", agent_endpoint="
Sources
Related stories

China Telecom's Xing4.0 Runs a 29B Coding Agent on 19GB Locally
Subtopic Mixture Of Experts · Long Context · Vision Language Takeaways − China Telecom released Xing4.0-29B-A4B , a 29B MoE with only 4B active parameters. Native 256K context extensible to 512K using MLA attention and 64 routed experts.

Red Hat Shrinks Nemotron 3.5 Lightning's 30B Agent Model by Half With FP8
Takeaways − Red Hat AI released an FP8 quantized build of NVIDIA Nemotron 3.5 Lightning 30B A3B. Cuts GPU memory and disk by roughly 50% versus the BF16 reference weights.

Cohere's North Small Translate Beats DeepL Across 50 Languages With Downloadable Weights
Subtopic Dpo · Fine Tuning · Distillation Takeaways − Cohere released North Small Translate , a 218B / 25B-active MoE translation model with open weights. Scores 83.60 on WMT26 across 50+ languages, beating DeepL NextGen (81.37) and Google Translate (68.20).

Add secure Web Search to Claude Desktop with Amazon Bedrock AgentCore
Claude Desktop on Amazon Bedrock provides powerful AI assistance, but without integrated web search, responses are limited to the model’s training knowledge…