Fine-tune a search agent with multi-turn RL on Amazon SageMaker AI | Amazon Web Services

https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/10/01/ML-21889-featured-image.png

Search agents powered by large language models (LLMs) are transforming how enterprises retrieve information. Rather than requiring users to craft the perfect query, a search agent autonomously decides what to search for, which retrieval strategy to use, and when to stop searching. It does this across multiple rounds of interaction, refining its approach based on what it has already retrieved.

However, getting this multi-step behavior to work well is hard. No base model arrives knowing your tools or your environment. Prompt a small model and you rarely get dependable multi-turn behavior. Prompt a frontier model and it often works, but you pay for that capability in latency and cost. Fine-tuning offers a third path: you teach a small model your tools and environment directly. The result is a small model’s speed and cost with the reliability that would otherwise require a frontier model.

Even though fine-tuning is the natural next...

Copyright of this story solely belongs to aws.amazon.com. To see the full text click HERE

Read more

http://www.techmeme.com/img/techmeme_sq328.png

Tavus unveils Griffin, the “first Human Interaction Model”, which it says passed the “video Turing test”, with 48% of users thinking it was human in live chats

Sponsor Posts Subquadratic: the LLM built for 12M-token reasoning — SubQ can reason across entire codebases and document sets in one pass with no RAG workarounds. Read how SubQ 1.1 Small holds near-perfect retrieval out to 12M tokens. Introducing Campus: The digital home for educational institutions — Every educational institution needs