Research finds AI agents haven't quite mastered real-world browsing tasks despite claiming they can

https://cdn.mos.cms.futurecdn.net/jwrMJ6cMNHe3jurU5dv9S7-1920-80.jpg
  • Not a single agent scored the full 20 out of 20
  • Claude for Chrome performed better than the ChatGPT Chrome Extension
  • With performance varying by testing category, Decodo advises selecting an agent based on planned usage

New Decodo research has criticized AI agents for still not being able to conduct real-world browsing tasks autonomously, including tasks like form filling, completing transactions, having cross-tab awareness and handling third-party integrations.

In fact, the study analyzed 45 AI agents across 10 different capabilities and found that not a single one could achieve the maximum score of 20.

The testing is also said to have exposed discrepancies between what vendors and AI developers say their agents are capable of, and what they can actually deliver on.

Agentic AI isn't at the level of autonomous browsing, yet

Claude for Chrome was the highest-scoring agent, reaching 18 points – two below the theoretical maximum. Its OpenAI...

Copyright of this story solely belongs to techradar.com. To see the full text click HERE

Read more